about 3 hours ago
Palo Alto, CA, USASenior
Base Salary
$195k - $262k/yr
Responsibilities
- Build and maintain distributed training infrastructure for various ML workloads.
- Integrate and extend frameworks such as Megatron-LM and DeepSpeed.
- Implement and debug parallelism strategies for model training.
- Build components for RL training including rollout and evaluation.
- Profile and improve GPU utilization and training throughput.
- Diagnose failures across multiple layers of the training infrastructure.
- Create reproducible training runs and operational tooling.
- Collaborate with research scientists to develop scalable systems.
- Write clear documentation including design and incident reports.
Requirements
- Strong Python and PyTorch engineering skills.
- Hands-on experience with distributed model training and large-scale ML systems.
- Practical understanding of transformer training bottlenecks.
- Experience debugging production training jobs across multiple GPUs.
- Ability to reason quantitatively about system performance metrics.
- Strong communication skills for collaboration with various teams.
Benefits
- 100% company-paid medical, dental, and vision coverage for employees and families.
- 401(k) plan with up to 4% company match and immediate vesting.
- 20 weeks paid parental leave for primary caregivers.
- Remote work reimbursement up to $85/month.
- Company-paid short-term, long-term, and life insurance coverage.
- Competitive compensation and career growth opportunities.