GrepJob
Nebius

ML Systems Engineer, Large-Scale Model Training & RL Infrastructure

Nebius
Apply
about 3 hours ago
Palo Alto, CA, USASenior

Base Salary

$195k - $262k/yr

Responsibilities

  • Build and maintain distributed training infrastructure for various ML workloads.
  • Integrate and extend frameworks such as Megatron-LM and DeepSpeed.
  • Implement and debug parallelism strategies for model training.
  • Build components for RL training including rollout and evaluation.
  • Profile and improve GPU utilization and training throughput.
  • Diagnose failures across multiple layers of the training infrastructure.
  • Create reproducible training runs and operational tooling.
  • Collaborate with research scientists to develop scalable systems.
  • Write clear documentation including design and incident reports.

Requirements

  • Strong Python and PyTorch engineering skills.
  • Hands-on experience with distributed model training and large-scale ML systems.
  • Practical understanding of transformer training bottlenecks.
  • Experience debugging production training jobs across multiple GPUs.
  • Ability to reason quantitatively about system performance metrics.
  • Strong communication skills for collaboration with various teams.

Benefits

  • 100% company-paid medical, dental, and vision coverage for employees and families.
  • 401(k) plan with up to 4% company match and immediate vesting.
  • 20 weeks paid parental leave for primary caregivers.
  • Remote work reimbursement up to $85/month.
  • Company-paid short-term, long-term, and life insurance coverage.
  • Competitive compensation and career growth opportunities.

Categories

AI & MLBackendData Engineering