8 months ago
Base Salary
$310k - $460k/yr
Responsibilities
- Design and build distributed failure detection, tracing, and profiling systems for large-scale AI training jobs.
- Develop tooling to identify slow, faulty, or misbehaving nodes and provide actionable system visibility.
- Improve observability, reliability, and performance across the training platform.
- Debug and resolve issues in complex, high-throughput distributed systems.
- Collaborate with systems, infrastructure, and research teams to evolve platform capabilities.
- Extend and adapt failure detection and tracing systems for new training paradigms and workloads.
Requirements
- Experience writing low-level software where system details matter.
- Understanding of hardware, operating systems, networking, concurrency, and distributed systems.
- Background in high-performance computing or low-level systems engineering.
- Interest in performance, stability, observability, large-scale debugging, and operational automation.
About OpenAI
OpenAI builds and deploys large-scale AI models and tools—including ChatGPT, GPT-4–class models, DALL·E, and Whisper—sold via APIs and enterprise subscriptions to developers and businesses. It monetizes through usage-based API pricing and ChatGPT Plus/Team/Enterprise, and also reaches customers via Microsoft’s Azure OpenAI Service. Founded in 2015 and headquartered in San Francisco, it operates as a private partnership.
