5 months ago
Base Salary
$295k - $555k/yr
Responsibilities
- Build and refine performance models that translate microbenchmark results into cost-to-serve estimates.
- Analyze inference workloads end to end across applications, models, and fleet infrastructure.
- Enhance tooling to identify bottlenecks affecting latency and throughput.
- Partner with other teams to turn performance insights into concrete improvements and project the effects of future changes on inference.
Requirements
- Deep expertise in performance profiling, benchmarking, analysis, and optimization.
- Ability to reason from first principles about distributed systems, model inference, and hardware efficiency.
- Comfort working across application behavior, kernels, accelerators, networking, and fleet scheduling.
- Interest in collaborating with engineering and research teams to improve real production systems.
Categories
About OpenAI
OpenAI builds and deploys large-scale AI models and tools—including ChatGPT, GPT-4–class models, DALL·E, and Whisper—sold via APIs and enterprise subscriptions to developers and businesses. It monetizes through usage-based API pricing and ChatGPT Plus/Team/Enterprise, and also reaches customers via Microsoft’s Azure OpenAI Service. Founded in 2015 and headquartered in San Francisco, it operates as a private partnership.
