OpenAI

Software Engineer, Inference - Performance Optimization

OpenAI
Apply
5 months ago

Base Salary

$295k - $555k/yr

Responsibilities

  • Build and refine performance models that translate microbenchmark results into cost-to-serve estimates.
  • Analyze inference workloads end to end across applications, models, and fleet infrastructure.
  • Enhance tooling to identify bottlenecks affecting latency and throughput.
  • Partner with other teams to turn performance insights into concrete improvements and project the effects of future changes on inference.

Requirements

  • Deep expertise in performance profiling, benchmarking, analysis, and optimization.
  • Ability to reason from first principles about distributed systems, model inference, and hardware efficiency.
  • Comfort working across application behavior, kernels, accelerators, networking, and fleet scheduling.
  • Interest in collaborating with engineering and research teams to improve real production systems.

Categories

OpenAI

About OpenAI

10,000+ employees

OpenAI builds and deploys large-scale AI models and tools—including ChatGPT, GPT-4–class models, DALL·E, and Whisper—sold via APIs and enterprise subscriptions to developers and businesses. It monetizes through usage-based API pricing and ChatGPT Plus/Team/Enterprise, and also reaches customers via Microsoft’s Azure OpenAI Service. Founded in 2015 and headquartered in San Francisco, it operates as a private partnership.

Contact me