Base Salary
$250k - $380k/yr
Responsibilities
- Design and maintain standardized dataset APIs, including APIs for multimodal data that cannot fit in memory.
- Build proactive testing and scale-validation pipelines for dataset loading at GPU scale.
- Integrate datasets into training and inference pipelines with teammates.
- Document and maintain discoverable, consistent dataset interfaces.
- Establish safeguards and validation systems to keep standardized datasets reproducible and unchanged.
- Debug performance bottlenecks in distributed dataset loading, including straggler systems.
- Provide visualization and inspection tools to identify dataset errors, bugs, and bottlenecks.
Requirements
- Strong engineering fundamentals with experience in distributed systems, data pipelines, or infrastructure.
- Experience building APIs, modular code, and scalable abstractions with attention to user experience.
- Comfort debugging performance bottlenecks across large fleets of machines.
- Bonus: background in data math, probability, or distributed data theory.
- Bonus: experience with GPU-scale distributed systems or dataset scaling for real-time data.
Categories
About OpenAI
OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. AI is an extremely powerful tool that must be created with safety and human needs at its core. OpenAI is dedicated to putting that alignment of interests first — ahead of profit. To achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. Our investment in diversity, equity, and inclusion is ongoing, executed through a wide range of initiatives, and championed and supported by leadership. At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.