Lambda

Senior Site Reliability Engineer - Managed Kubernetes

Lambda
Apply
3 hours ago
Bellevue, WA, USA +2 moreSenior
H1B Sponsor

Base Salary

$240k - $356k/yr

Responsibilities

  • Operate and maintain bare-metal Kubernetes clusters scaling to thousands of nodes
  • Handle cluster degradation, recovery, resizing, and critical incident response using fleet management tools
  • Participate in an on-call rotation for critical incidents
  • Assist customers with Kubernetes questions, workload integration, storage, and authentication
  • Collaborate with HPC Ops and Datacenter Ops teams on low-level and cross-functional issues
  • Create Python and Go tooling to automate platform-quality validation
  • Design, build, and maintain scalable Kubernetes control plane services, operators, and custom controllers
  • Automate Kubernetes cluster provisioning, upgrades, patching, and deletion
  • Define and implement SLOs and SLIs for Kubernetes services, workloads, and platform reliability

Requirements

  • 6+ years of experience in SRE, operations engineering, or a similar role
  • Deep knowledge of running Linux clusters and systems
  • Strong programming skills in Go and Python
  • Experience with GitOps, ArgoCD, Helm, and Kubernetes operators
  • Production experience operating Kubernetes clusters in on-premises, EKS, GKE, or similar environments
  • Ability to work independently with limited direction and collaboratively as part of a team
  • Ability to work with customers during incidents through tickets, live messaging, or larger calls
  • Familiarity with Prometheus, Grafana, Fluent Bit, and CI/CD pipelines
  • Experience provisioning Kubernetes with kubeadm, Cluster API, or similar tools
  • Deep Kubernetes expertise including CRDs, CSI, CNI, and Kubernetes operator coding
  • Exposure to HPC clusters, AI/ML workloads, or large-scale GPU clusters
  • Experience with hybrid or multicloud Kubernetes environments
  • Contributions to CNCF projects or Kubernetes SIGs are a plus

Benefits

  • Hybrid work arrangement requiring presence in the San Francisco, San Jose, or Bellevue office 4 days per week, with Tuesday designated as the work-from-home day
  • Health, dental, and vision coverage for employees and dependents
  • Wellness and commuter stipends for select roles
  • 401(k) plan with a 2% company match for U.S. employees
  • Flexible paid time off plan
  • Cash and equity compensation

Tech Stack

GoGrafanaHelmKubernetesLinuxPrometheusPython

Categories

DevOpsSite Reliability
Lambda

About Lambda

501-1,000 employees

The Superintelligence Cloud