7 months ago
London, United KingdomMid Level
Responsibilities
- Design, build, and maintain software that orchestrates and monitors machine learning workloads on large supercomputers.
- Profile and optimize the software stack for computation orchestration at frontier scale.
- Improve reliability, observability, and fault tolerance for long-running jobs.
- Debug complex distributed systems issues across large clusters.
- Adapt the training runtime to changing machine learning systems needs and support researchers.
- Work across the Python and Rust stack.
Requirements
- Experience developing distributed systems rather than only operating them.
- Proficiency in Python and Rust or another systems programming language such as C++.
- Solid Linux knowledge and comfort with systems-level debugging, performance analysis, and memory profiling.
- Experience developing asynchronous and concurrent systems.
- Strong software engineering skills with a focus on performance, correctness, and reliability.
Benefits
- Hybrid work model with 3 days per week in the London office.
- Relocation assistance is available to new employees.
About OpenAI
OpenAI builds and deploys large-scale AI models and tools—including ChatGPT, GPT-4–class models, DALL·E, and Whisper—sold via APIs and enterprise subscriptions to developers and businesses. It monetizes through usage-based API pricing and ChatGPT Plus/Team/Enterprise, and also reaches customers via Microsoft’s Azure OpenAI Service. Founded in 2015 and headquartered in San Francisco, it operates as a private partnership.
