2 months ago
Base Salary
$266k - $445k/yr
Responsibilities
- Design and implement the LLM inference runtime for frontier models on custom silicon.
- Build scheduling, continuous batching, memory management, KV-cache management, and execution orchestration.
- Develop distributed execution strategies across chips, hosts, and racks, including model partitioning, communication, and synchronization.
- Optimize latency, throughput, memory efficiency, and hardware utilization across model architectures and serving workloads.
- Partner with kernel, compiler, architecture, and silicon teams to co-design interfaces and resolve performance bottlenecks.
- Enable new model features, execution patterns, numerical formats, and hardware capabilities in a reliable production runtime.
- Create profiling, observability, benchmarking, and performance-modeling tools.
- Debug correctness, performance, and reliability issues across model code, runtime software, communication layers, and hardware.
- Translate workload insights into requirements for future silicon and system architecture.
Requirements
- Strong systems programming experience in C++, Rust, Python, or comparable performance-oriented environments.
- Experience building or optimizing runtimes, distributed systems, compilers, kernels, model-serving infrastructure, or adjacent systems software.
- Understanding of modern LLM inference, including prefill and decode behavior, batching, KV-cache tradeoffs, and model parallelism.
- Ability to reason quantitatively about latency, throughput, compute intensity, memory bandwidth, communication, and utilization.
- Experience profiling and debugging performance across multiple layers of a hardware-software stack.
- Ability to design clean abstractions while retaining low-level control for specialized hardware performance.
- Ability to collaborate across model, systems, compiler, kernel, and hardware teams on ambiguous technical problems.
- Commitment to production quality, including correctness, observability, reliability, maintainability, and graceful behavior at scale.
- Candidates may need to meet legal status requirements under U.S. export control laws and regulations.
Categories
About OpenAI
OpenAI builds and deploys large-scale AI models and tools—including ChatGPT, GPT-4–class models, DALL·E, and Whisper—sold via APIs and enterprise subscriptions to developers and businesses. It monetizes through usage-based API pricing and ChatGPT Plus/Team/Enterprise, and also reaches customers via Microsoft’s Azure OpenAI Service. Founded in 2015 and headquartered in San Francisco, it operates as a private partnership.
