Base Salary
$295k - $555k/yr
Responsibilities
- Design and implement inference infrastructure for large-scale multimodal models.
- Optimize systems for high-throughput, low-latency delivery of image and audio inputs and outputs.
- Transition experimental research workflows into reliable production services.
- Collaborate with researchers, infrastructure teams, and product engineers to deploy advanced model capabilities.
- Improve GPU utilization, tensor parallelism, and hardware abstraction layers.
Requirements
- Experience building and scaling inference systems for LLMs or multimodal models.
- Experience with GPU-based machine learning workloads and the performance characteristics of large models, particularly for image or audio data.
- Comfort working across networking, distributed compute, and high-throughput data handling.
- Familiarity with inference tooling such as vLLM, TensorRT-LLM, or custom model-parallel systems.
- Ability to own problems end-to-end in experimental and fast-moving environments.
- Experience with image generation or audio synthesis models in production is preferred.
- Exposure to distributed ML training or system-efficient model design is preferred.
Categories
About OpenAI
OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. AI is an extremely powerful tool that must be created with safety and human needs at its core. OpenAI is dedicated to putting that alignment of interests first — ahead of profit. To achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. Our investment in diversity, equity, and inclusion is ongoing, executed through a wide range of initiatives, and championed and supported by leadership. At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.