OpenAI

Software Engineer, Data Infrastructure - Research

OpenAI
Apply
12 months ago

Base Salary

$250k - $380k/yr

Responsibilities

  • Design and maintain standardized dataset APIs, including APIs for multimodal data that cannot fit in memory.
  • Build proactive testing and scale-validation pipelines for dataset loading at GPU scale.
  • Integrate datasets into training and inference pipelines with teammates.
  • Document and maintain discoverable, consistent dataset interfaces.
  • Establish safeguards and validation systems to keep standardized datasets reproducible and unchanged.
  • Debug performance bottlenecks in distributed dataset loading, including straggler systems.
  • Provide visualization and inspection tools to identify dataset errors, bugs, and bottlenecks.

Requirements

  • Strong engineering fundamentals with experience in distributed systems, data pipelines, or infrastructure.
  • Experience building APIs, modular code, and scalable abstractions with attention to user experience.
  • Comfort debugging performance bottlenecks across large fleets of machines.
  • Bonus: background in data math, probability, or distributed data theory.
  • Bonus: experience with GPU-scale distributed systems or dataset scaling for real-time data.

Categories

BackendData Engineering
OpenAI

About OpenAI

10,000+ employees

OpenAI builds and deploys large-scale AI models and tools—including ChatGPT, GPT-4–class models, DALL·E, and Whisper—sold via APIs and enterprise subscriptions to developers and businesses. It monetizes through usage-based API pricing and ChatGPT Plus/Team/Enterprise, and also reaches customers via Microsoft’s Azure OpenAI Service. Founded in 2015 and headquartered in San Francisco, it operates as a private partnership.

Contact me