Base Salary
$230k - $490k/yr
Responsibilities
- Develop and optimize high-performance cluster systems for compute-intensive AI workloads
- Ensure the reliability, scalability, and efficiency of cluster infrastructure
- Design, build, bring up, operate, and maintain large-scale compute clusters
- Write software for cluster orchestration, resource allocation, and lifecycle automation
- Collaborate with researchers and engineering teams to understand compute needs and optimize resource allocation
- Implement and uphold security measures across cluster systems
Requirements
- Experience designing scalable, reliable, and secure compute clusters using distributed systems principles
- Strong programming skills in Python, Go, or similar languages
- Experience working in public cloud environments, especially Azure
- Familiarity with high-performance computing, GPU workloads, or AI/ML compute patterns
- Ability to take initiative and build effectively in a fast-paced, dynamic environment
Benefits
- Hybrid work model with 3 days per week in the San Francisco office
- Relocation assistance for new employees
About OpenAI
OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. AI is an extremely powerful tool that must be created with safety and human needs at its core. OpenAI is dedicated to putting that alignment of interests first — ahead of profit. To achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. Our investment in diversity, equity, and inclusion is ongoing, executed through a wide range of initiatives, and championed and supported by leadership. At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.