15 hours ago
Base Salary
$177k - $327k/yr
Responsibilities
- Define the model CI strategy for frontier models running on custom silicon and integrated accelerator systems.
- Build production CI/CD pipelines for regression testing, performance benchmarking, software stability, and release qualification.
- Design test matrices covering models, configurations, hardware generations, software components, and deployment environments.
- Create performance baselines, regression detection, bisect and triage workflows, and failure ownership processes.
- Establish GitOps practices for reproducible configuration, promotion, rollback, auditability, and environment consistency.
- Shape monorepo architecture, dependency management, build and test boundaries, change validation, and developer workflows.
- Develop orchestration, artifact management, caching, scheduling, observability, and capacity controls for accelerator-backed CI.
- Partner with model, compiler, kernel, runtime, firmware, validation, and hardware teams to translate release risks into automated gates.
- Improve CI reliability, speed, debuggability, and cost efficiency while maintaining production confidence.
Requirements
- Experience building or operating large-scale CI/CD, developer infrastructure, test automation, or production engineering systems.
- Understanding of GitOps, release promotion, rollback, reproducibility, and policy-driven automation.
- Experience designing monorepos, build systems, dependency graphs, test selection, or change-impact analysis.
- Ability to design regression and performance-benchmarking systems with stable baselines, useful metrics, and actionable failure diagnosis.
- Proficiency in Python, Go, Rust, C++, or another language used for infrastructure and automation.
- Experience with distributed systems, schedulers, containers, clusters, or heterogeneous compute infrastructure.
- Ability to collaborate with model and systems engineers to make accelerator workloads repeatable for production qualification.
- Strong focus on reliability, observability, developer experience, and reducing time from code change to trustworthy signal.
About OpenAI
OpenAI builds and deploys large-scale AI models and tools—including ChatGPT, GPT-4–class models, DALL·E, and Whisper—sold via APIs and enterprise subscriptions to developers and businesses. It monetizes through usage-based API pricing and ChatGPT Plus/Team/Enterprise, and also reaches customers via Microsoft’s Azure OpenAI Service. Founded in 2015 and headquartered in San Francisco, it operates as a private partnership.
