OpenAI

Software Engineer, Fleet Infrastructure

OpenAI
Apply
over 1 year ago
San Francisco, CA, USAMid Level
H1B sponsor

Base Salary

$230k - $490k/yr

Responsibilities

  • Design, implement, and operate GPU compute fleet components including job scheduling, cluster management, snapshot delivery, and CI/CD systems
  • Build scheduling and quota systems that maximize GPU utilization
  • Develop automation for Kubernetes cluster provisioning and upgrades
  • Support research workflows with service frameworks and deployment systems
  • Improve model startup times through high-performance snapshot delivery from blob storage to hardware caching
  • Work with researchers and product teams to understand workload requirements
  • Collaborate with hardware, infrastructure, and business teams to deliver highly utilized and reliable services

Requirements

  • Experience with hyperscale compute systems
  • Strong programming skills
  • Experience working in public clouds, especially Azure
  • Experience working with Kubernetes
  • An execution-focused mindset and rigorous attention to user requirements
  • Understanding of AI/ML workloads is a bonus

Benefits

  • Hybrid work model requiring 3 days per week in the San Francisco office
  • Relocation assistance is available to new employees

Tech Stack

Categories

OpenAI

About OpenAI

10,000+ employees

OpenAI builds and deploys large-scale AI models and tools—including ChatGPT, GPT-4–class models, DALL·E, and Whisper—sold via APIs and enterprise subscriptions to developers and businesses. It monetizes through usage-based API pricing and ChatGPT Plus/Team/Enterprise, and also reaches customers via Microsoft’s Azure OpenAI Service. Founded in 2015 and headquartered in San Francisco, it operates as a private partnership.

Contact me