
Research Engineer - RL Infrastructure
Prime Intellect2 months ago
Base Salary
$150k - $350k/yr
Responsibilities
- Build and optimize systems infrastructure for large-scale RL and distributed training workloads in the prime-rl framework.
- Improve end-to-end training efficiency across compute, memory, networking, and scheduling layers.
- Implement low-level optimizations for kernels, communication paths, and runtimes.
- Develop distributed training systems using data, tensor, and pipeline parallelism.
- Help shape asynchronous rollout and post-training architecture.
- Contribute to open-source libraries and internal infrastructure for frontier-scale model training.
- Collaborate with researchers and infrastructure engineers to turn bottlenecks into systems improvements.
Requirements
- Strong systems engineering experience in AI/ML infrastructure, particularly large-scale model training or inference.
- Deep familiarity with PyTorch and distributed training frameworks such as PyTorch Distributed, DeepSpeed, FSDP, Megatron, vLLM, or Ray.
- Experience optimizing kernels, memory movement, communication overhead, or parallelization strategies.
- Hands-on experience with data, tensor, and pipeline parallelism.
- Strong understanding of GPU architecture, profiling, and performance debugging.
- Ability to identify system bottlenecks and drive improvements from first principles.
- Experience with CUDA or Triton kernels, compiler/runtime optimization, RL training infrastructure, multi-node GPU clusters, high-performance networking, or open-source ML systems is especially valued.
Benefits
- Cash compensation of $150-350k plus equity.
- Flexible work arrangements with remote or in-person work from the San Francisco office.
- Visa sponsorship and relocation support for international candidates.
- Quarterly team offsites, hackathons, conferences, and learning opportunities.
- Work with a highly technical, high-agency team building open-superintelligence infrastructure.