5 months ago
Base Salary
$230k - $385k/yr
Responsibilities
- Improve development workflows for engineers working on model performance infrastructure.
- Design and improve CI/CD, release, validation, and testing pipelines.
- Build and maintain tools that improve reliability, iteration speed, and engineering confidence.
- Partner with engineers to identify friction in testing, debugging, deployment, and development workflows.
- Contribute to infrastructure supporting performance-critical training and inference systems.
- Improve developer experience across Python-heavy codebases and performance-oriented infrastructure.
- Contribute to the Triton project and related performance engineering systems.
Requirements
- Strong experience with CI/CD, developer infrastructure, testing systems, tooling, or build and release workflows.
- Strong Python skills and interest in building reliable, scalable developer tools and infrastructure.
- Experience improving large-scale engineering workflows, especially CI reliability, test infrastructure, and debugging velocity.
- Experience with the PyTorch ecosystem is highly relevant.
- Experience with C++ or Rust is a nice-to-have but not required.
- Ability to work autonomously, collaborate closely with technical teams, and operate effectively in an ambiguous environment.
- Direct inference or model performance experience is not required, but enthusiasm for learning the domain is expected.
About OpenAI
OpenAI builds and deploys large-scale AI models and tools—including ChatGPT, GPT-4–class models, DALL·E, and Whisper—sold via APIs and enterprise subscriptions to developers and businesses. It monetizes through usage-based API pricing and ChatGPT Plus/Team/Enterprise, and also reaches customers via Microsoft’s Azure OpenAI Service. Founded in 2015 and headquartered in San Francisco, it operates as a private partnership.
