5 days ago
Base Salary
$165k - $225k/yr
Responsibilities
- Train large byte-native and multimodal foundation models across massive, heterogeneous corpora.
- Implement and evaluate model architectures, training objectives, and optimization methods.
- Develop stable pre-training recipes and conduct scaling experiments for novel architectures.
- Run ablations and analyze training dynamics, model behavior, and base-model quality.
- Collaborate with data and distributed training engineers to improve training efficiency, reliability, and scalability.
- Write robust, performant, production-grade training code and infrastructure.
Requirements
- 5+ years of experience in machine learning research or engineering, including developing and pre-training large language or multimodal foundation models.
- Strong general software engineering skills and the ability to write robust, performant training code.
- Solid understanding of deep learning fundamentals, modern pre-training methods, and related literature.
- Ability to implement research ideas and evaluate them using baselines, ablations, metrics, and analysis.
- Hands-on experience running GPU-based training workloads and familiarity with distributed training.
- MS in Computer Science, Machine Learning, Artificial Intelligence, Mathematics, or a related field.
- Preferred qualifications include a PhD, extensive JAX/Flax/XLA experience, multi-node pre-training with FSDP, ZeRO, or Megatron, experience developing training recipes and scaling experiments, and ownership of monitored, reproducible training and evaluation pipelines.
Benefits
- Medical, dental, and vision insurance.
- 401k plan.
- Daily lunch, snacks, and beverages.
- Flexible time off.
- Competitive salary and equity.
Categories
AI ResearchML Engineering
