
Research Engineer - Distributed Training
Prime Intellect2 months ago
Base Salary
$150k - $350k/yr
Responsibilities
- Build and optimize distributed training infrastructure for pre-training and large-scale reinforcement learning workloads through the prime-rl framework.
- Improve end-to-end training efficiency across compute, memory, networking, and scheduling layers.
- Design and implement low-level optimizations involving kernels, communication paths, and runtimes.
- Develop distributed training systems using data, tensor, and pipeline parallelism.
- Help shape RL training architecture, including asynchronous rollout and post-training systems.
- Contribute to open-source libraries and internal infrastructure for frontier-scale model training.
- Collaborate with researchers and infrastructure engineers to turn bottlenecks into systems improvements.
Requirements
- Strong systems engineering experience in AI/ML infrastructure, particularly large-scale model training or inference.
- Deep familiarity with PyTorch and distributed training frameworks such as PyTorch Distributed, DeepSpeed, FSDP, Megatron, vLLM, or Ray.
- Experience optimizing training performance across kernels, memory movement, communication overhead, or parallelization strategies.
- Hands-on experience with data parallelism, tensor parallelism, and pipeline parallelism.
- Strong understanding of GPU architecture, profiling, and performance debugging.
- Ability to identify system bottlenecks across the stack and drive improvements from first principles.
- Comfort working in a fast-moving environment with ambiguous problems and high ownership.
- Experience with CUDA or Triton kernel optimization, ML compiler or runtime optimization, RL training infrastructure, multi-node GPU clusters, high-performance networking, or open-source ML systems is especially valued.
- Interest in publishing technical work or sharing engineering insights through technical writing is especially valued.
Benefits
- Cash compensation of $150-350k plus equity incentives.
- Flexible work arrangements with the option to work remotely or in person at San Francisco offices.
- Visa sponsorship and relocation assistance for international candidates.
- Quarterly team off-sites, hackathons, conferences, and learning opportunities.
- Opportunity to work with a mission-driven team focused on frontier AI and infrastructure.