2 months ago
Base Salary
$267k - $356k/yr
Responsibilities
- Own the reliability, performance, and capacity health of Lambda’s production storage fleet across all data centers.
- Build and maintain monitoring, dashboards, and alerting for storage performance, capacity, and hardware failures.
- Investigate and resolve storage incidents using telemetry, logs, and performance profiling.
- Automate ticketing, escalation, and incident-response workflows.
- Design and maintain self-healing automation for drive replacement, node swaps, rebuild monitoring, and capacity rebalancing.
- Implement CI/CD pipelines for storage automation and tooling.
- Automate deployment and configuration of software-defined storage across new and existing sites with Storage Engineering, Fleet Orchestration, and Release Engineering.
- Diagnose low-level storage, I/O, hardware, and networking issues with hardware and networking teams.
- Participate in an on-call rotation focused on reducing MTTR and recurring alerts.
Requirements
- 5+ years of experience operating Linux systems in production or HPC environments, with hands-on storage experience at scale on scale-out or software-defined platforms such as CEPH, Lustre, or GPFS.
- Hands-on experience operating software-defined storage platforms at scale and integrating with their management and data-plane APIs.
- Experience owning production storage incidents from initial alert through root cause analysis and postmortem.
- Experience with monitoring and logging platforms such as Prometheus, Grafana, Alertmanager, Datadog, or SumoLogic, including dashboards and alert or pager routing.
- Experience with Kubernetes and GitOps tooling such as ArgoCD, Helm, or Kustomize, including troubleshooting.
- Experience with CI/CD tooling such as GitHub Actions, Jenkins, or BuildKite; containerization with Docker or Podman; and systems programming in Python or Go.
- Experience with Infrastructure as Code using Terraform or Ansible.
- Understanding of file, object, block, and structured storage protocols and technologies including NFS, SMB, S3, NVMe-oF/TCP, vector databases, and SQL.
- Preferred experience with VAST, Weka, NetApp, Dell PowerScale, GPFS, or Lustre.
- Preferred experience writing or operating Kubernetes CSI drivers.
- Preferred experience with SR-IOV, KVM, QEMU, GPUDirect Storage, RDMA, InfiniBand, RoCE, ethtool, mlxlink, and clush.
- Open-source storage project contributions are a plus.
Benefits
- Hybrid schedule requiring presence in the San Francisco or San Jose office 4 days per week, with Tuesday designated as the work-from-home day.
- Health, dental, and vision coverage for employees and dependents.
- Wellness and commuter stipends for select roles.
- 401(k) plan with a 2% company match for U.S. employees.
- Flexible paid time off plan.
- Cash and equity compensation are offered, with no amounts specified.
Tech Stack
Categories
DevOpsSite Reliability
About Lambda
Lambda provides GPU cloud computing and on-prem AI hardware—servers, clusters, and workstations—for teams training and serving large ML models. Its products include NVIDIA H100/A100 instances, managed clusters, and the Lambda Stack software, sold via usage-based cloud pricing and hardware sales. Founded in 2012 and headquartered in San Francisco, the privately held company serves researchers, startups, enterprises, and hyperscalers.
