over 1 year ago
Base Salary
$230k - $490k/yr
Responsibilities
- Build and maintain automation systems for provisioning and managing server fleets.
- Develop tools to monitor server health, performance, and lifecycle events.
- Collaborate with cluster, networking, and infrastructure teams.
- Partner with external operators to maintain high quality.
- Identify and fix performance bottlenecks and inefficiencies.
- Continuously improve automation to reduce manual work.
Requirements
- Experience managing large-scale server environments.
- Strength in both building systems and operationalizing them.
- Proficiency in Python, Go, or similar languages.
- Strong knowledge of Linux, networking, and server hardware.
- Comfort analyzing noisy data with SQL, PromQL, Pandas, or comparable tools.
- Prior hardware expertise is not required.
- Bonus: knowledge of PCIe, InfiniBand, networking, power management, kernel performance tuning, IPMI, Redfish, HPC, distributed systems, hardware development, Prometheus, or Grafana.
About OpenAI
OpenAI builds and deploys large-scale AI models and tools—including ChatGPT, GPT-4–class models, DALL·E, and Whisper—sold via APIs and enterprise subscriptions to developers and businesses. It monetizes through usage-based API pricing and ChatGPT Plus/Team/Enterprise, and also reaches customers via Microsoft’s Azure OpenAI Service. Founded in 2015 and headquartered in San Francisco, it operates as a private partnership.
