2 months ago
Base Salary
$240k - $356k/yr
Responsibilities
- Operate and scale critical platform services across Lambda’s physical data centers.
- Improve the reliability of compute provisioning, instance lifecycle, and regional orchestration systems.
- Build monitoring, alerting, and tracing for service health, provisioning latency, and customer-impacting failures.
- Define SLIs, SLOs, error budgets, and operational readiness standards.
- Automate detection and remediation of configuration drift, failed workflows, and orphaned resources.
- Build safe deployment, rollback, and disaster recovery workflows using infrastructure as code and GitOps.
- Design fault-isolation mechanisms to reduce blast radius and prevent cascading failures.
- Lead production incident response, postmortems, and durable corrective actions.
- Partner with Compute, Networking, Storage, Security, and Support teams.
- Participate in on-call and improve its sustainability through automation and better tooling.
- Mentor engineers and raise the organization’s reliability standards.
Requirements
- 7+ years of experience in site reliability, infrastructure, distributed systems, or production software engineering.
- Deep experience operating Kubernetes in production, including its architecture, scheduling, networking, resource management, upgrades, and failure modes.
- Experience with physical data centers, private cloud, hybrid cloud, or environments without full reliance on managed services.
- Proficiency with Terraform or similar infrastructure-as-code tools.
- Experience building CI/CD or GitOps workflows with tools such as Argo CD, Flux, Helm, or Kustomize.
- Experience with observability platforms such as OpenTelemetry, Prometheus, Grafana, or Datadog.
- Ability to build production-quality tooling in Go, Python, or a similar language.
- Understanding of distributed systems concepts including consistency, retries, idempotency, backpressure, and partial failure.
- Experience defining and operating against SLIs and SLOs and leading high-severity incidents.
- Preferred experience includes AI infrastructure, GPU platforms, high-performance computing, multi-region distributed systems, Kubernetes controllers and operators, CRDs, admission control, scheduler extensions, etcd operations, Linux systems, container runtimes, cgroups, storage, networking, chaos engineering, Kubernetes RBAC, OIDC, workload identity, certificate management, and compliance frameworks such as SOC 2 or ISO 27001.
Benefits
- Hybrid work arrangement requiring presence in the San Francisco, San Jose, or Bellevue office 4 days per week, with Tuesday designated as the work-from-home day.
- Health, dental, and vision coverage for employees and dependents.
- Wellness and commuter stipends for select roles.
- 401(k) plan with a 2% company match for U.S. employees.
- Flexible paid time off plan.
- Cash and equity compensation are offered, with no specific amounts stated.
Tech Stack
Categories
DevOpsSite Reliability
About Lambda
Lambda provides GPU cloud computing and on-prem AI hardware—servers, clusters, and workstations—for teams training and serving large ML models. Its products include NVIDIA H100/A100 instances, managed clusters, and the Lambda Stack software, sold via usage-based cloud pricing and hardware sales. Founded in 2012 and headquartered in San Francisco, the privately held company serves researchers, startups, enterprises, and hyperscalers.
