GrepJob
Lambda

Senior Site Reliability Engineer - Storage

Lambda
Apply
about 4 hours ago
San Francisco, CA, USA or San Jose, CA, USASenior
H1B Sponsor

Base Salary

$267k - $356k/yr

Responsibilities

  • Own reliability, performance, and capacity health for Lambda’s production storage fleet across all data centers.
  • Build and maintain monitoring, dashboards, and alerting for storage performance, capacity, and hardware failures.
  • Investigate and resolve storage incidents using telemetry, logs, and performance profiling.
  • Automate ticketing, escalation, incident response, drive replacement, node swaps, rebuild monitoring, and capacity rebalancing.
  • Implement CI/CD pipelines for storage automation and tooling.
  • Automate deployment and configuration of software-defined storage across new and existing sites.
  • Diagnose low-level storage, I/O, networking, NIC, RDMA/RoCE, InfiniBand, and multipathing issues with hardware and networking teams.
  • Participate in an on-call rotation and reduce mean time to resolution through automation.

Requirements

  • At least 5 years of experience operating Linux systems in production or HPC environments, with hands-on storage experience at scale.
  • Experience operating software-defined storage platforms at scale and integrating with management and data-plane APIs.
  • Ability to own production storage incidents from alert through root cause analysis and postmortem.
  • Experience with monitoring and logging platforms, including dashboard creation and alert or pager routing.
  • Experience with Kubernetes, GitOps tooling, and hands-on troubleshooting.
  • Experience with CI/CD tooling, containerization, and systems programming in Python or Go.
  • Experience with Infrastructure as Code using Terraform or Ansible.
  • Understanding of file, object, block, and structured storage protocols and technologies.
  • Preferred experience with VAST, Weka, NetApp, Dell PowerScale, GPFS, Lustre, Kubernetes CSI drivers, SR-IOV, KVM/QEMU, GPUDirect Storage, RDMA, InfiniBand, RoCE, ethtool, mlxlink, clush, or open-source storage projects.

Benefits

  • Presence in the San Francisco or San Jose office 4 days per week, with Tuesday designated as the work-from-home day.
  • Generous cash and equity compensation.
  • Health, dental, and vision coverage for employees and dependents.
  • Wellness and commuter stipends for select roles.
  • 401(k) plan with a 2% company match for USA employees.
  • Flexible paid time off plan.

Tech Stack

AnsibleBuildkiteDatadogDockerGitHub ActionsGoGrafanaHelmJenkinsKubernetesLinuxPrometheusPythonTerraform

Categories

Lambda

About Lambda

501-1,000 employees

The Superintelligence Cloud