over 1 year ago
Base Salary
$295k - $380k/yr
Responsibilities
- Prototype and enable OpenAI’s AI software stack on new and exploratory accelerator platforms.
- Optimize large-scale model performance and distributed AI workloads across diverse hardware environments.
- Develop kernels, sharding mechanisms, and system scaling strategies for emerging accelerators.
- Collaborate on model-level and low-level optimizations, including PyTorch-based workloads.
- Model system performance, debug bottlenecks, and drive end-to-end optimization.
- Work with hardware teams and vendors to evaluate alternative platforms and adapt the software stack to their architectures.
- Contribute to runtime improvements, compute and communication overlapping, and scaling for frontier AI workloads.
Requirements
- At least 3 years of experience in AI infrastructure, including kernels, systems, or hardware-software co-design.
- Hands-on experience with data-center-scale AI accelerator platforms such as TPUs, custom silicon, or exploratory architectures.
- Strong understanding of kernels, sharding, runtime systems, or distributed scaling techniques.
- Experience optimizing LLMs, CNNs, or recommender models for hardware efficiency.
- Experience with performance modeling, system debugging, and adapting software stacks to novel architectures.
- Exposure to mobile accelerators is welcome, while data-center-scale AI hardware experience is preferred.
- Ability to work across multiple levels of the stack, rapidly prototype solutions, and navigate ambiguity during early hardware bring-up.
About OpenAI
OpenAI builds and deploys large-scale AI models and tools—including ChatGPT, GPT-4–class models, DALL·E, and Whisper—sold via APIs and enterprise subscriptions to developers and businesses. It monetizes through usage-based API pricing and ChatGPT Plus/Team/Enterprise, and also reaches customers via Microsoft’s Azure OpenAI Service. Founded in 2015 and headquartered in San Francisco, it operates as a private partnership.
