5 months ago
Base Salary
$230k - $405k/yr
Responsibilities
- Provision, bootstrap, scale, and manage the lifecycle of large Kubernetes clusters.
- Build software abstractions that unify multiple clusters and provide seamless interfaces to training workloads.
- Own bare-metal node bring-up, including firmware upgrades, for repeatable large-scale deployment.
- Improve cluster restart times and accelerate firmware and operating system upgrade cycles.
- Integrate networking and hardware health systems across servers, switches, and data-center infrastructure.
- Develop monitoring and observability systems to detect issues early and maintain cluster stability under extreme load.
- Improve training performance, infrastructure utilization, resilience, cost efficiency, and uptime across large compute fleets.
Requirements
- Experience as an infrastructure, systems, or distributed systems engineer in large-scale or high-availability environments.
- Strong knowledge of Kubernetes internals, cluster scaling patterns, and containerized workloads.
- Proficiency in compute infrastructure concepts, including compute, networking, storage, and security, with experience automating cluster or data-center operations.
- Background with GPU workloads, firmware management, or high-performance computing is a bonus.
Tech Stack
Categories
About OpenAI
OpenAI builds and deploys large-scale AI models and tools—including ChatGPT, GPT-4–class models, DALL·E, and Whisper—sold via APIs and enterprise subscriptions to developers and businesses. It monetizes through usage-based API pricing and ChatGPT Plus/Team/Enterprise, and also reaches customers via Microsoft’s Azure OpenAI Service. Founded in 2015 and headquartered in San Francisco, it operates as a private partnership.
