3 months ago
London, United KingdomSenior
Responsibilities
- Design, build, and operate software managing large-scale GPU infrastructure supporting ChatGPT inference.
- Build internal platforms, tooling, and AI-powered agents that automate fleet operations and reduce operational overhead.
- Improve observability, reliability, and operational efficiency across thousands of GPUs.
- Develop systems for capacity planning, scheduling, fleet health monitoring, and incident response.
- Identify infrastructure bottlenecks and implement solutions that improve utilization, scalability, and performance.
- Partner with research, platform, networking, and systems teams to improve the compute platform.
- Establish engineering best practices around operational excellence, automation, and infrastructure reliability.
Requirements
- 5+ years of software engineering experience building production infrastructure.
- Strong programming skills in Go, Python, C++, Rust, or similar systems languages.
- Experience designing and operating highly available distributed systems.
- Experience with GPU infrastructure, high-performance computing, ML infrastructure, or large-scale compute platforms.
- Experience with Kubernetes, cloud infrastructure, Linux, networking, and observability tooling.
- Strong debugging, systems design, and operational problem-solving skills.
- Strong communication skills and experience collaborating across engineering organizations.
About OpenAI
OpenAI builds and deploys large-scale AI models and tools—including ChatGPT, GPT-4–class models, DALL·E, and Whisper—sold via APIs and enterprise subscriptions to developers and businesses. It monetizes through usage-based API pricing and ChatGPT Plus/Team/Enterprise, and also reaches customers via Microsoft’s Azure OpenAI Service. Founded in 2015 and headquartered in San Francisco, it operates as a private partnership.
