Lambda

Senior Site Reliability Engineer

Lambda
Apply
3 hours ago
Bellevue, WA, USA or San Francisco, CA, USASenior
H1B Sponsor

Base Salary

$240k - $356k/yr

Responsibilities

  • Build and operate monitoring and alerting for fabric, GPU, power, thermal, and job-level cluster health signals.
  • Remotely deploy and configure large-scale HPC clusters for AI workloads using automation.
  • Automate cluster lifecycle management for operating systems, firmware, drivers, and networking with Ansible and Terraform.
  • Create runbooks and automated remediations for common cluster failure modes for Support and HPC Support teams.
  • Troubleshoot cluster issues involving InfiniBand/RoCE, NCCL, GPU-direct, fabric, switching, and power in coordination with on-site deployment teams.
  • Participate in on-call rotations and lead incident response for cluster-level problems.
  • Maintain Standard Operating Procedures and communicate requirements to engineering teams to improve simplification, stability, and operational efficiency.

Requirements

  • 7+ years of experience in Site Reliability Engineering, HPC Engineering, DevOps, or a similar role.
  • Strong understanding of modern AI infrastructure, GPU architectures, and hardware performance optimization.
  • Strong understanding of Linux-based systems in distributed environments.
  • Experience configuring and troubleshooting InfiniBand, RoCE, CLOS fabrics, 100GbE, Ethernet and switching, GPU-direct, and NCCL environments.
  • Solid understanding of Python and Go, with experience collaborating with software engineering teams to improve internal tooling.
  • Experience with monitoring and alerting tools such as Prometheus, Grafana, and ClickHouse.
  • Proficiency with automation and configuration management tools such as Ansible and Terraform.
  • Experience with machine learning or deep learning frameworks such as PyTorch and TensorFlow is preferred.
  • Experience with benchmarking tools such as DeepSpeed and MLPerf is preferred.
  • Knowledge of Docker and Kubernetes is preferred.
  • Experience building or operating HPC resources is preferred.
  • Depth in the NVIDIA hardware and firmware ecosystem is preferred.
  • Experience with data center power and thermal design is preferred.
  • Background in chaos engineering or similar reliability testing methodologies is preferred.
  • Understanding of compliance frameworks such as SOC 2 and ISO 27001 is preferred.

Benefits

  • Requires working from the San Francisco or Bellevue office four days per week, with Tuesday designated as the work-from-home day.
  • Health, dental, and vision coverage for employees and dependents.
  • Wellness and commuter stipends for select roles.
  • 401(k) plan with a 2% company match for U.S. employees.
  • Flexible paid time off plan.
  • Cash and equity compensation are offered, with no specific amounts stated.

Tech Stack

AnsibleClickHouseDockerGoGrafanaKubernetesLinuxPrometheusPythonPyTorchTensorFlowTerraform

Categories

DevOpsSite Reliability
Lambda

About Lambda

501-1,000 employees

The Superintelligence Cloud