Lambda

Senior Site Reliability Engineer - Storage

Lambda
Apply
2 months ago
San Francisco, CA, USA or San Jose, CA, USASenior
H1B sponsor

Base Salary

$267k - $356k/yr

Responsibilities

  • Own the reliability, performance, and capacity health of Lambda’s production storage fleet across all data centers.
  • Build and maintain monitoring, dashboards, and alerting for storage performance, capacity, and hardware failures.
  • Investigate and resolve storage incidents using telemetry, logs, and performance profiling.
  • Automate ticketing, escalation, and incident-response workflows.
  • Design and maintain self-healing automation for drive replacement, node swaps, rebuild monitoring, and capacity rebalancing.
  • Implement CI/CD pipelines for storage automation and tooling.
  • Automate deployment and configuration of software-defined storage across new and existing sites with Storage Engineering, Fleet Orchestration, and Release Engineering.
  • Diagnose low-level storage, I/O, hardware, and networking issues with hardware and networking teams.
  • Participate in an on-call rotation focused on reducing MTTR and recurring alerts.

Requirements

  • 5+ years of experience operating Linux systems in production or HPC environments, with hands-on storage experience at scale on scale-out or software-defined platforms such as CEPH, Lustre, or GPFS.
  • Hands-on experience operating software-defined storage platforms at scale and integrating with their management and data-plane APIs.
  • Experience owning production storage incidents from initial alert through root cause analysis and postmortem.
  • Experience with monitoring and logging platforms such as Prometheus, Grafana, Alertmanager, Datadog, or SumoLogic, including dashboards and alert or pager routing.
  • Experience with Kubernetes and GitOps tooling such as ArgoCD, Helm, or Kustomize, including troubleshooting.
  • Experience with CI/CD tooling such as GitHub Actions, Jenkins, or BuildKite; containerization with Docker or Podman; and systems programming in Python or Go.
  • Experience with Infrastructure as Code using Terraform or Ansible.
  • Understanding of file, object, block, and structured storage protocols and technologies including NFS, SMB, S3, NVMe-oF/TCP, vector databases, and SQL.
  • Preferred experience with VAST, Weka, NetApp, Dell PowerScale, GPFS, or Lustre.
  • Preferred experience writing or operating Kubernetes CSI drivers.
  • Preferred experience with SR-IOV, KVM, QEMU, GPUDirect Storage, RDMA, InfiniBand, RoCE, ethtool, mlxlink, and clush.
  • Open-source storage project contributions are a plus.

Benefits

  • Hybrid schedule requiring presence in the San Francisco or San Jose office 4 days per week, with Tuesday designated as the work-from-home day.
  • Health, dental, and vision coverage for employees and dependents.
  • Wellness and commuter stipends for select roles.
  • 401(k) plan with a 2% company match for U.S. employees.
  • Flexible paid time off plan.
  • Cash and equity compensation are offered, with no amounts specified.

Tech Stack

AnsibleBuildkiteDatadogDockerGitHub ActionsGoGrafanaHelmJenkinsKubernetesLinuxPrometheusPythonSQLTerraform

Categories

DevOpsSite Reliability
Lambda

About Lambda

501-1,000 employees

Lambda provides GPU cloud computing and on-prem AI hardware—servers, clusters, and workstations—for teams training and serving large ML models. Its products include NVIDIA H100/A100 instances, managed clusters, and the Lambda Stack software, sold via usage-based cloud pricing and hardware sales. Founded in 2012 and headquartered in San Francisco, the privately held company serves researchers, startups, enterprises, and hyperscalers.

Contact me