over 1 year ago
Base Salary
$295k - $555k/yr
Responsibilities
- Design and implement inference infrastructure for large-scale multimodal models.
- Optimize systems for high-throughput, low-latency delivery of image and audio inputs and outputs.
- Transition experimental research workflows into reliable production services.
- Collaborate with researchers, infrastructure teams, and product engineers to deploy advanced model capabilities.
- Improve GPU utilization, tensor parallelism, and hardware abstraction layers.
Requirements
- Experience building and scaling inference systems for LLMs or multimodal models.
- Experience with GPU-based machine learning workloads and the performance characteristics of large models, particularly for image or audio data.
- Comfort working across networking, distributed compute, and high-throughput data handling.
- Familiarity with inference tooling such as vLLM, TensorRT-LLM, or custom model-parallel systems.
- Ability to own problems end-to-end in experimental and fast-moving environments.
- Experience with image generation or audio synthesis models in production is preferred.
- Exposure to distributed ML training or system-efficient model design is preferred.
Categories
About OpenAI
OpenAI builds and deploys large-scale AI models and tools—including ChatGPT, GPT-4–class models, DALL·E, and Whisper—sold via APIs and enterprise subscriptions to developers and businesses. It monetizes through usage-based API pricing and ChatGPT Plus/Team/Enterprise, and also reaches customers via Microsoft’s Azure OpenAI Service. Founded in 2015 and headquartered in San Francisco, it operates as a private partnership.
