over 2 years ago
Base Salary
$380k - $555k/yr
Responsibilities
- Design and implement efficient collective operations in C++ and CUDA in collaboration with ML researchers.
- Ensure large training jobs fully utilize the network transports available in the supercomputers.
- Develop simulations to inform future supercomputer network designs.
Requirements
- Background in low-level, performance-critical software.
- Experience writing distributed algorithms using RDMA.
- Comfort writing low-level, performance-sensitive CPU and/or GPU code.
- Familiarity with network simulation techniques.
- Experience with collective communication is a bonus.
Benefits
- Hybrid work model with 3 days in the San Francisco office per week
- Relocation assistance for new employees
About OpenAI
OpenAI builds and deploys large-scale AI models and tools—including ChatGPT, GPT-4–class models, DALL·E, and Whisper—sold via APIs and enterprise subscriptions to developers and businesses. It monetizes through usage-based API pricing and ChatGPT Plus/Team/Enterprise, and also reaches customers via Microsoft’s Azure OpenAI Service. Founded in 2015 and headquartered in San Francisco, it operates as a private partnership.
