12 months ago
Base Salary
$250k - $445k/yr
Responsibilities
- Profile end-to-end training runs to identify compute, communication, and storage bottlenecks.
- Optimize GPU utilization and throughput for large-scale distributed model training.
- Improve kernel efficiency, scheduling, and collective communication performance with runtime and systems engineers.
- Implement model graph transforms to improve end-to-end throughput.
- Build tooling to monitor and visualize MFU, throughput, and uptime across clusters.
- Partner with researchers to ensure new model architectures scale efficiently during pre-training.
- Contribute to infrastructure decisions that improve the reliability and efficiency of large training jobs.
Requirements
- Strong programming skills in Python and C++; Rust or CUDA experience is a plus.
- Experience running distributed training jobs on multi-GPU systems or HPC clusters.
- Exposure to PyTorch, JAX, or TensorFlow and understanding of large-scale training loops.
- Ability to debug complex distributed systems, rigorously measure efficiency, and translate profiling data into engineering improvements.
- Familiarity with NCCL, MPI, or UCX is preferred.
- Experience with large-scale data loading and checkpointing systems is preferred.
- Prior work on training runtime, distributed scheduling, or ML compiler optimization is preferred.
Benefits
- Hybrid work model with three days in the San Francisco office per week
- Relocation assistance for new employees
About OpenAI
OpenAI builds and deploys large-scale AI models and tools—including ChatGPT, GPT-4–class models, DALL·E, and Whisper—sold via APIs and enterprise subscriptions to developers and businesses. It monetizes through usage-based API pricing and ChatGPT Plus/Team/Enterprise, and also reaches customers via Microsoft’s Azure OpenAI Service. Founded in 2015 and headquartered in San Francisco, it operates as a private partnership.
