3 hours ago
Base Salary
$240k - $356k/yr
Responsibilities
- Build and operate monitoring and alerting for fabric, GPU, power, thermal, and job-level cluster health signals.
- Remotely deploy and configure large-scale HPC clusters for AI workloads using automation.
- Automate cluster lifecycle management for operating systems, firmware, drivers, and networking with Ansible and Terraform.
- Create runbooks and automated remediations for common cluster failure modes for Support and HPC Support teams.
- Troubleshoot cluster issues involving InfiniBand/RoCE, NCCL, GPU-direct, fabric, switching, and power in coordination with on-site deployment teams.
- Participate in on-call rotations and lead incident response for cluster-level problems.
- Maintain Standard Operating Procedures and communicate requirements to engineering teams to improve simplification, stability, and operational efficiency.
Requirements
- 7+ years of experience in Site Reliability Engineering, HPC Engineering, DevOps, or a similar role.
- Strong understanding of modern AI infrastructure, GPU architectures, and hardware performance optimization.
- Strong understanding of Linux-based systems in distributed environments.
- Experience configuring and troubleshooting InfiniBand, RoCE, CLOS fabrics, 100GbE, Ethernet and switching, GPU-direct, and NCCL environments.
- Solid understanding of Python and Go, with experience collaborating with software engineering teams to improve internal tooling.
- Experience with monitoring and alerting tools such as Prometheus, Grafana, and ClickHouse.
- Proficiency with automation and configuration management tools such as Ansible and Terraform.
- Experience with machine learning or deep learning frameworks such as PyTorch and TensorFlow is preferred.
- Experience with benchmarking tools such as DeepSpeed and MLPerf is preferred.
- Knowledge of Docker and Kubernetes is preferred.
- Experience building or operating HPC resources is preferred.
- Depth in the NVIDIA hardware and firmware ecosystem is preferred.
- Experience with data center power and thermal design is preferred.
- Background in chaos engineering or similar reliability testing methodologies is preferred.
- Understanding of compliance frameworks such as SOC 2 and ISO 27001 is preferred.
Benefits
- Requires working from the San Francisco or Bellevue office four days per week, with Tuesday designated as the work-from-home day.
- Health, dental, and vision coverage for employees and dependents.
- Wellness and commuter stipends for select roles.
- 401(k) plan with a 2% company match for U.S. employees.
- Flexible paid time off plan.
- Cash and equity compensation are offered, with no specific amounts stated.
Categories
DevOpsSite Reliability
