over 1 year ago
Base Salary
$230k - $490k/yr
Responsibilities
- Design, implement, and operate GPU compute fleet components including job scheduling, cluster management, snapshot delivery, and CI/CD systems
- Build scheduling and quota systems that maximize GPU utilization
- Develop automation for Kubernetes cluster provisioning and upgrades
- Support research workflows with service frameworks and deployment systems
- Improve model startup times through high-performance snapshot delivery from blob storage to hardware caching
- Work with researchers and product teams to understand workload requirements
- Collaborate with hardware, infrastructure, and business teams to deliver highly utilized and reliable services
Requirements
- Experience with hyperscale compute systems
- Strong programming skills
- Experience working in public clouds, especially Azure
- Experience working with Kubernetes
- An execution-focused mindset and rigorous attention to user requirements
- Understanding of AI/ML workloads is a bonus
Benefits
- Hybrid work model requiring 3 days per week in the San Francisco office
- Relocation assistance is available to new employees
Tech Stack
Categories
About OpenAI
OpenAI builds and deploys large-scale AI models and tools—including ChatGPT, GPT-4–class models, DALL·E, and Whisper—sold via APIs and enterprise subscriptions to developers and businesses. It monetizes through usage-based API pricing and ChatGPT Plus/Team/Enterprise, and also reaches customers via Microsoft’s Azure OpenAI Service. Founded in 2015 and headquartered in San Francisco, it operates as a private partnership.
