almost 2 years ago
Base Salary
$230k - $490k/yr
Responsibilities
- Develop and optimize high-performance cluster systems for compute-intensive AI workloads
- Ensure the reliability, scalability, and efficiency of cluster infrastructure
- Design, build, bring up, operate, and maintain large-scale compute clusters
- Write software for cluster orchestration, resource allocation, and lifecycle automation
- Collaborate with researchers and engineering teams to understand compute needs and optimize resource allocation
- Implement and uphold security measures across cluster systems
Requirements
- Experience designing scalable, reliable, and secure compute clusters using distributed systems principles
- Strong programming skills in Python, Go, or similar languages
- Experience working in public cloud environments, especially Azure
- Familiarity with high-performance computing, GPU workloads, or AI/ML compute patterns
- Ability to take initiative and build effectively in a fast-paced, dynamic environment
Benefits
- Hybrid work model with 3 days per week in the San Francisco office
- Relocation assistance for new employees
About OpenAI
OpenAI builds and deploys large-scale AI models and tools—including ChatGPT, GPT-4–class models, DALL·E, and Whisper—sold via APIs and enterprise subscriptions to developers and businesses. It monetizes through usage-based API pricing and ChatGPT Plus/Team/Enterprise, and also reaches customers via Microsoft’s Azure OpenAI Service. Founded in 2015 and headquartered in San Francisco, it operates as a private partnership.
