
ML Systems Engineer
Periodic Labs4 months ago
Base Salary
$300k - $400k/yr
Responsibilities
- Build rack- and topology-aware scheduling for GB-series GPUs across Ray, Slurm, and Kubernetes.
- Build online and offline profilers to identify bottlenecks across training and inference systems.
- Implement direct S3 checkpoint streaming to reduce I/O bottlenecks in large-scale training.
- Benchmark RL training configurations across model sizes, batch strategies, and hardware topologies.
- Write and optimize communication and GPU kernels for maximum hardware throughput.
- Design zero-copy RDMA weight synchronization between training and inference.
- Build fast sandbox execution environments for model-generated actions and reward feedback.
- Engage with the SGLang, Megatron, and Ray communities through upstream contributions and roadmap influence.
- Collaborate with RL and pretraining researchers to co-design algorithms and infrastructure.
Requirements
- Bachelor’s degree or an equivalent combination of education, training, or experience.
- Experience with large-scale inference infrastructure, including load balancing, traffic shifting, scheduling, and production serving architecture.
- Experience with low-level systems programming involving RDMA, NVLink, kernel-level work, or network stack optimization.
- Experience with GPU cluster scheduling and orchestration across Ray, Slurm, or Kubernetes, including rack topology and hardware locality.
- Experience writing or optimizing CUDA kernels, communication primitives, or distributed training collective operations.
- Experience profiling and benchmarking distributed ML systems across compute, memory, and network bottlenecks.
- Experience managing and streaming checkpoints at scale, including direct cloud storage integration.
- Experience building or contributing to open-source ML infrastructure projects such as SGLang, Megatron-LM, vLLM, or Ray.
- Experience collaborating directly with ML researchers on algorithm-infrastructure co-design.
Benefits
- The lab is located in Menlo Park, with preference for candidates in Menlo Park or San Francisco and flexibility based on the role.
- Visa sponsorship is available, with legal support for the process.
Tech Stack
Categories
About Periodic Labs
We're building AI scientists and the autonomous laboratories for them to operate. Join us: https://jobs.ashbyhq.com/periodic-labs