OpenAI

Software Engineer, Workload Enablement

OpenAI
Apply
6 months ago
Seattle, WA, USA or San Francisco, CA, USASenior
H1B sponsor

Base Salary

$293k - $455k/yr

Responsibilities

  • Port and validate inference and training workloads on new platforms and SKUs, driving correctness, performance, and stability to an internal readiness standard.
  • Build end-to-end benchmarks and stress tests covering CPU, GPU, memory, frontend, scale-up and scale-out networking, WAN traffic, NVLink, RDMA collectives, storage, and thermals.
  • Analyze distributed training and inference performance, including collective tuning, compute and communication overlap, kernel bottlenecks, memory bandwidth, and scheduling effects.
  • Create repeatable test harnesses for CI and lab environments that generate pass/fail results, performance scores, and regression detection.
  • Partner with systems and fleet bring-up engineers on containerization, Kubernetes integration, telemetry hooks, failure triage, operational usability, and scalability.
  • Work with vendors and internal stakeholders to produce bug reports, minimal reproductions, and prioritized issue lists.

Requirements

  • Bachelor’s degree in computer science or electrical engineering, or equivalent practical experience.
  • At least five years of experience in ML systems, performance engineering, distributed systems, or HPC.
  • Hands-on experience with PyTorch and modern LLM training and inference stacks.
  • Understanding of large-scale distributed training, including data, model, and pipeline parallelism and collective communications.
  • Experience with RDMA and debugging or optimizing NCCL or RCCL communications libraries and their hardware and network interactions.
  • Proficiency in Python and ability to read or write performance-critical code; C++, CUDA, or HIP experience is a plus.
  • Strong profiling and debugging skills using tools such as Nsight, rocprof, perf, and flamegraphs.
  • Experience building workload-shaped benchmarks and stress or fault tests that reflect production behavior.
  • Familiarity with RDMA networking, transport tuning, network topology, and congestion effects on collectives.
  • Experience validating workloads in Kubernetes and turning research code into robust, repeatable infrastructure.
  • Hands-on lab experience with early hardware such as new network interface cards, GPUs, accelerators, or racks.
OpenAI

About OpenAI

10,000+ employees

OpenAI builds and deploys large-scale AI models and tools—including ChatGPT, GPT-4–class models, DALL·E, and Whisper—sold via APIs and enterprise subscriptions to developers and businesses. It monetizes through usage-based API pricing and ChatGPT Plus/Team/Enterprise, and also reaches customers via Microsoft’s Azure OpenAI Service. Founded in 2015 and headquartered in San Francisco, it operates as a private partnership.

Contact me