5 months ago
Base Salary
$230k - $385k/yr
Responsibilities
- Improve deploy-gate tooling and infrastructure to validate inference engine releases for correctness, performance, and regression freedom.
- Strengthen release, validation, branching, deployment, canary, asynchronous, and large-scale validation workflows.
- Harden CI, testing, and validation infrastructure so failures are actionable and trustworthy.
- Reduce flaky and noisy failures caused by infrastructure instability, GPU scheduling, or test-environment issues.
- Build automation for failure triage, ownership detection, debugging, and escalation.
- Improve observability, rollout safety, release automation, and developer self-service tooling.
- Partner with inference, research developer productivity, engine acceleration, and infrastructure teams to reduce developer friction.
Requirements
- Experience with CI/CD systems, testing infrastructure, release tooling, developer productivity, or large-scale build and validation systems.
- Comfort working in Python-heavy environments and debugging complex distributed systems.
- Ability to build automation for triage, debugging, ownership detection, and operational effectiveness.
- Strong ownership, developer empathy, and comfort navigating ambiguous cross-functional problems.
- C++ experience is helpful for inference engine code, CI build issues, or performance-sensitive systems, but is not required.
- Prior inference experience is not required.
About OpenAI
OpenAI builds and deploys large-scale AI models and tools—including ChatGPT, GPT-4–class models, DALL·E, and Whisper—sold via APIs and enterprise subscriptions to developers and businesses. It monetizes through usage-based API pricing and ChatGPT Plus/Team/Enterprise, and also reaches customers via Microsoft’s Azure OpenAI Service. Founded in 2015 and headquartered in San Francisco, it operates as a private partnership.
