2 months ago
Palo Alto, CA, USASenior
Base Salary
$195k - $262k/yr
Responsibilities
- Build and maintain distributed training infrastructure for SFT, continued pretraining, preference optimization, and reinforcement-learning workloads.
- Integrate and extend large-scale training and RL frameworks and internal systems.
- Implement and debug tensor, pipeline, sequence/context, expert, and data parallelism strategies.
- Build rollout, reward-model serving, replay/data buffer, checkpointing, evaluation, and experiment-orchestration components.
- Profile and improve GPU utilization, communication efficiency, memory usage, and training throughput.
- Diagnose failures across NCCL, CUDA, PyTorch, Ray, schedulers, storage, networking, and checkpointing layers.
- Create reproducible training runs, launch scripts, dashboards, runbooks, and operational tooling for research users.
- Partner with research scientists to convert algorithmic training recipes into scalable and debuggable systems.
- Write design documents, incident reports, benchmark reports, and operating guides.
Requirements
- Strong Python and PyTorch engineering skills.
- Hands-on experience with distributed model training, large-scale ML systems, or GPU cluster workloads.
- Practical understanding of transformer training bottlenecks, memory pressure, gradient and optimizer state, communication overhead, and checkpointing.
- Experience debugging production or research training jobs across multiple GPUs or nodes.
- Ability to reason quantitatively about throughput, utilization, memory, reliability, cost, and research velocity.
- Strong communication and collaboration skills with researchers, ML engineers, platform engineers, and leadership.
- Experience with Megatron-LM, DeepSpeed, PyTorch FSDP/DTensor, Ray, Slurm, Kubernetes, or large internal training platforms is preferred.
- Experience with reinforcement-learning infrastructure frameworks such as verl, slime, AReaL, OpenRLHF, TRL, or custom PPO, GRPO, or RLHF systems is preferred.
- Familiarity with NCCL, CUDA, Triton, Nsight, InfiniBand, RDMA, RoCE, H100, H200, B200 clusters, or storage and network bottlenecks is preferred.
- Experience supporting SFT, DPO, PPO, GRPO, RLAIF, reward-model serving, rollout generation, or agent-training workloads is preferred.
- Open-source contributions to distributed training, RL infrastructure, PyTorch, Ray, Megatron, DeepSpeed, or related systems are preferred.
Benefits
- 100% company-paid medical, dental, and vision coverage for employees and families.
- 401(k) plan with up to 4% company match and immediate vesting.
- 20 weeks of paid parental leave for primary caregivers and 12 weeks for secondary caregivers.
- Remote work reimbursement of up to $85 per month for mobile and internet.
- Company-paid short-term disability, long-term disability, and life insurance.
- Remote work is supported, with applicants required to be authorized to work in the country where they apply.
Tech Stack
Categories
About Nebius
The Nebius AI Cloud brings powerful full-stack infrastructure for AI developers and practitioners across startups, enterprises and science institutes to build and deploy generative AI applications and rapidly deliver scientific breakthroughs by training and running ML models within a secure, high-performance, and cost-optimized cloud environment.
