10 hours ago
Base Salary
$266k - $500k/yr
Responsibilities
- Build and operate infrastructure for large-scale model training and evaluation.
- Develop shared inference and grading platforms with automated capacity management, health monitoring, and performance visibility.
- Improve compute scheduling and resource allocation to reduce idle GPU time and enable rapid workload recovery.
- Diagnose bottlenecks across training, inference, and orchestration and improve end-to-end performance.
- Build self-service tools, automated validation, and observability for researchers launching experiments, diagnosing issues, and comparing results.
- Own projects from bottleneck identification and solution design through deployment and operation.
Requirements
- Strong software engineering fundamentals and experience building or operating large-scale distributed systems.
- Experience with ML infrastructure, inference systems, GPU performance, or infrastructure tooling.
- Ability to take ownership of open-ended problems and debug across system boundaries.
- Interest in building infrastructure that enables personal AGI and large-scale AI research.
- Ability to use measurements to guide performance and reliability improvements.
Categories
About OpenAI
OpenAI builds and deploys large-scale AI models and tools—including ChatGPT, GPT-4–class models, DALL·E, and Whisper—sold via APIs and enterprise subscriptions to developers and businesses. It monetizes through usage-based API pricing and ChatGPT Plus/Team/Enterprise, and also reaches customers via Microsoft’s Azure OpenAI Service. Founded in 2015 and headquartered in San Francisco, it operates as a private partnership.
