Prime Intellect

Research Engineer - Distributed Training

Prime Intellect
Apply
2 months ago
Remote, United States or San Francisco, CA, USASenior
H1B sponsor

Base Salary

$150k - $350k/yr

Responsibilities

  • Build and optimize distributed training infrastructure for pre-training and large-scale reinforcement learning workloads through the prime-rl framework.
  • Improve end-to-end training efficiency across compute, memory, networking, and scheduling layers.
  • Design and implement low-level optimizations involving kernels, communication paths, and runtimes.
  • Develop distributed training systems using data, tensor, and pipeline parallelism.
  • Help shape RL training architecture, including asynchronous rollout and post-training systems.
  • Contribute to open-source libraries and internal infrastructure for frontier-scale model training.
  • Collaborate with researchers and infrastructure engineers to turn bottlenecks into systems improvements.

Requirements

  • Strong systems engineering experience in AI/ML infrastructure, particularly large-scale model training or inference.
  • Deep familiarity with PyTorch and distributed training frameworks such as PyTorch Distributed, DeepSpeed, FSDP, Megatron, vLLM, or Ray.
  • Experience optimizing training performance across kernels, memory movement, communication overhead, or parallelization strategies.
  • Hands-on experience with data parallelism, tensor parallelism, and pipeline parallelism.
  • Strong understanding of GPU architecture, profiling, and performance debugging.
  • Ability to identify system bottlenecks across the stack and drive improvements from first principles.
  • Comfort working in a fast-moving environment with ambiguous problems and high ownership.
  • Experience with CUDA or Triton kernel optimization, ML compiler or runtime optimization, RL training infrastructure, multi-node GPU clusters, high-performance networking, or open-source ML systems is especially valued.
  • Interest in publishing technical work or sharing engineering insights through technical writing is especially valued.

Benefits

  • Cash compensation of $150-350k plus equity incentives.
  • Flexible work arrangements with the option to work remotely or in person at San Francisco offices.
  • Visa sponsorship and relocation assistance for international candidates.
  • Quarterly team off-sites, hackathons, conferences, and learning opportunities.
  • Opportunity to work with a mission-driven team focused on frontier AI and infrastructure.

Tech Stack

Prime Intellect

About Prime Intellect

51-200 employees
Contact me