OpenAI

Software Engineer, GPU Infrastructure - HPC

OpenAI
Apply
over 1 year ago

Base Salary

$230k - $490k/yr

Responsibilities

  • Build and maintain automation systems for provisioning and managing server fleets.
  • Develop tools to monitor server health, performance, and lifecycle events.
  • Collaborate with cluster, networking, and infrastructure teams.
  • Partner with external operators to maintain high quality.
  • Identify and fix performance bottlenecks and inefficiencies.
  • Continuously improve automation to reduce manual work.

Requirements

  • Experience managing large-scale server environments.
  • Strength in both building systems and operationalizing them.
  • Proficiency in Python, Go, or similar languages.
  • Strong knowledge of Linux, networking, and server hardware.
  • Comfort analyzing noisy data with SQL, PromQL, Pandas, or comparable tools.
  • Prior hardware expertise is not required.
  • Bonus: knowledge of PCIe, InfiniBand, networking, power management, kernel performance tuning, IPMI, Redfish, HPC, distributed systems, hardware development, Prometheus, or Grafana.

Tech Stack

GoGrafanaLinuxPandasPrometheusPythonSQL

Categories

DevOpsSite Reliability
OpenAI

About OpenAI

10,000+ employees

OpenAI builds and deploys large-scale AI models and tools—including ChatGPT, GPT-4–class models, DALL·E, and Whisper—sold via APIs and enterprise subscriptions to developers and businesses. It monetizes through usage-based API pricing and ChatGPT Plus/Team/Enterprise, and also reaches customers via Microsoft’s Azure OpenAI Service. Founded in 2015 and headquartered in San Francisco, it operates as a private partnership.

Contact me