OpenAI

Software Engineer, Compute Infrastructure

OpenAI
Apply
5 months ago
London, United Kingdom +3 moreSenior
H1B sponsor

Base Salary

$230k - $405k/yr

Responsibilities

  • Provision, bootstrap, scale, and manage the lifecycle of large Kubernetes clusters.
  • Build software abstractions that unify multiple clusters and provide seamless interfaces to training workloads.
  • Own bare-metal node bring-up, including firmware upgrades, for repeatable large-scale deployment.
  • Improve cluster restart times and accelerate firmware and operating system upgrade cycles.
  • Integrate networking and hardware health systems across servers, switches, and data-center infrastructure.
  • Develop monitoring and observability systems to detect issues early and maintain cluster stability under extreme load.
  • Improve training performance, infrastructure utilization, resilience, cost efficiency, and uptime across large compute fleets.

Requirements

  • Experience as an infrastructure, systems, or distributed systems engineer in large-scale or high-availability environments.
  • Strong knowledge of Kubernetes internals, cluster scaling patterns, and containerized workloads.
  • Proficiency in compute infrastructure concepts, including compute, networking, storage, and security, with experience automating cluster or data-center operations.
  • Background with GPU workloads, firmware management, or high-performance computing is a bonus.

Tech Stack

Categories

OpenAI

About OpenAI

10,000+ employees

OpenAI builds and deploys large-scale AI models and tools—including ChatGPT, GPT-4–class models, DALL·E, and Whisper—sold via APIs and enterprise subscriptions to developers and businesses. It monetizes through usage-based API pricing and ChatGPT Plus/Team/Enterprise, and also reaches customers via Microsoft’s Azure OpenAI Service. Founded in 2015 and headquartered in San Francisco, it operates as a private partnership.

Contact me