3 hours ago
Base Salary
$255k - $490k/yr
Responsibilities
- Design, build, and operate Kubernetes-based controllers and distributed services coordinating infrastructure across sites.
- Define APIs and resource models for consistent machine and cluster lifecycle operations across hardware platforms and providers.
- Build provisioning and configuration services integrating network boot, hardware management interfaces, firmware, operating-system images, drivers, and host configuration.
- Develop lifecycle management for discovery, allocation, provisioning, upgrades, maintenance, recovery, and decommissioning.
- Design reconciliation and recovery for concurrency, interrupted operations, partial failures, and staged rollouts.
- Improve control-plane throughput, API latency, and infrastructure convergence while respecting site and provider limits.
- Integrate new sites and GPU hardware generations into the platform in partnership with hardware, networking, data-center, and infrastructure teams.
Requirements
- Strong software engineering fundamentals and experience designing, implementing, and owning production distributed systems or infrastructure services.
- Experience developing infrastructure systems using Kubernetes APIs and reconciliation to manage resources.
- Understanding of bare-metal node provisioning, with depth in one or more of PXE, DHCP/DNS, baseboard management controllers, firmware, Linux, drivers, images, or configuration management.
- Ability to design reliable APIs and asynchronous workflows involving concurrency, consistency, idempotency, and failures across service and provider boundaries.
- Ability to diagnose reliability and performance problems across service, operating-system, and machine boundaries.
- Ability to collaborate across engineering specialties and communicate system behavior and technical tradeoffs clearly.
- Experience building multi-site or multi-region infrastructure control planes is preferred.
- Experience with GPU or HPC infrastructure and integrating multiple hardware platforms or infrastructure providers is preferred.
Tech Stack
About OpenAI
OpenAI builds and deploys large-scale AI models and tools—including ChatGPT, GPT-4–class models, DALL·E, and Whisper—sold via APIs and enterprise subscriptions to developers and businesses. It monetizes through usage-based API pricing and ChatGPT Plus/Team/Enterprise, and also reaches customers via Microsoft’s Azure OpenAI Service. Founded in 2015 and headquartered in San Francisco, it operates as a private partnership.
