Lambda

Staff Software Engineer - Managed Kubernetes

Lambda
Apply
28 days ago
Bellevue, WA, USA +2 moreStaff+
H1B Sponsor

Base Salary

$314k - $465k/yr

Responsibilities

  • Drive the technical vision and development of Lambda’s bare-metal Managed Kubernetes platform, including control plane scalability, multi-tenancy, cluster lifecycle management, and high availability.
  • Integrate and extend NVIDIA’s GPU orchestration and networking ecosystem, including GPU Operator, Network Operator, DCGM, NCCL, AICR, and Topograph.
  • Design GPU-aware orchestration, scheduling, inference platform services, model serving infrastructure, autoscaling, and multi-model deployment patterns.
  • Build the foundation for Managed Slurm on Kubernetes and support HPC workloads alongside Kubernetes workloads.
  • Define networking and storage requirements for AI workloads, including CNI integration, high-performance fabrics, RDMA, and GPUDirect.
  • Design self-healing systems, incident-response automation, root-cause analysis workflows, platform resilience, and chaos engineering programs.
  • Establish operational excellence through upgrade automation, security patching, and zero-downtime maintenance.
  • Set technical direction, lead design reviews, influence infrastructure roadmaps, mentor engineers, and standardize practices across teams.
  • Partner with Network, Storage, Security, Customer Success, NVIDIA, customers, and the open-source community.
  • Shape Lambda’s AIOps vision for capacity planning, anomaly detection, and predictive infrastructure maintenance.
  • Represent Lambda through technical writing, conference talks, and strategic customer engagements.

Requirements

  • 10+ years of experience in software engineering, platform engineering, or SRE, including at least 5 years focused on Kubernetes at scale.
  • Expert understanding of Kubernetes internals, including API machinery, controllers, schedulers, operators, CRDs, CSI, CNI, and extension patterns.
  • Holistic infrastructure expertise spanning compute, networking, storage, and security.
  • Strong production software engineering skills in Go and Python.
  • Deep experience with GPU orchestration in Kubernetes, including NVIDIA GPU Operator, device plugins, DCGM, MIG, time-slicing, and GPU-aware scheduling.
  • Proven technical leadership experience driving cross-team decisions, mentoring engineers, and influencing infrastructure direction.
  • Deep experience designing and operating managed services or multi-tenant platforms for external customers.
  • Strong understanding of distributed systems, including consensus, fault tolerance, consistency models, and graceful degradation.
  • Experience with observability at scale, including Prometheus, Grafana, distributed tracing, and actionable alerting.
  • Knowledge of Linux systems, L2-L7 networking, RDMA, InfiniBand, and RoCE.
  • Experience with infrastructure-as-code and GitOps workflows.
  • Preferred experience building or operating managed Kubernetes services such as GKE, EKS, or AKS, or working on Kubernetes control plane components.
  • Preferred experience with NVIDIA Network Operator, NCCL tuning, Topograph, AICR, or similar projects.
  • Familiarity with Slurm, KAI, Volcano, or Kueue and traditional or Kubernetes-native batch scheduling.
  • Background in confidential computing, ML infrastructure, customer migrations, or security and compliance in multi-tenant environments.
  • Familiarity with RBAC, Pod Security Standards, network policies, workload isolation, CNCF projects, Kubernetes SIGs, or NVIDIA open-source projects.

Benefits

  • Requires working from the San Francisco, San Jose, or Bellevue office 4 days per week, with Tuesday designated as the work-from-home day.
  • Health, dental, and vision coverage for employees and dependents.
  • Wellness and commuter stipends for select roles.
  • 401(k) plan with company match for U.S. employees.
  • Flexible paid time off plan.
  • Cash and equity compensation are offered.

Tech Stack

GoGrafanaKubernetesLinuxPrometheusPython

Categories

Lambda

About Lambda

501-1,000 employees

The Superintelligence Cloud

Contact me