OpenAI

Training: Process Management Engineer

OpenAI
Apply
7 months ago
London, United KingdomMid Level

Responsibilities

  • Design, build, and maintain software that orchestrates and monitors machine learning workloads on large supercomputers.
  • Profile and optimize the software stack for computation orchestration at frontier scale.
  • Improve reliability, observability, and fault tolerance for long-running jobs.
  • Debug complex distributed systems issues across large clusters.
  • Adapt the training runtime to changing machine learning systems needs and support researchers.
  • Work across the Python and Rust stack.

Requirements

  • Experience developing distributed systems rather than only operating them.
  • Proficiency in Python and Rust or another systems programming language such as C++.
  • Solid Linux knowledge and comfort with systems-level debugging, performance analysis, and memory profiling.
  • Experience developing asynchronous and concurrent systems.
  • Strong software engineering skills with a focus on performance, correctness, and reliability.

Benefits

  • Hybrid work model with 3 days per week in the London office.
  • Relocation assistance is available to new employees.
OpenAI

About OpenAI

10,000+ employees

OpenAI builds and deploys large-scale AI models and tools—including ChatGPT, GPT-4–class models, DALL·E, and Whisper—sold via APIs and enterprise subscriptions to developers and businesses. It monetizes through usage-based API pricing and ChatGPT Plus/Team/Enterprise, and also reaches customers via Microsoft’s Azure OpenAI Service. Founded in 2015 and headquartered in San Francisco, it operates as a private partnership.

Contact me