Lambda

Staff Software Engineer - Managed Kubernetes

Lambda
Apply
4 days ago

Base Salary

$314k - $465k/yr

Responsibilities

  • Drive the technical vision and development of Lambda’s bare-metal Managed Kubernetes platform, including control-plane scalability, multi-tenancy, cluster lifecycle management, and high availability.
  • Integrate and extend NVIDIA’s GPU orchestration ecosystem and design GPU-aware scheduling and placement systems.
  • Lead development of managed orchestration services, including Managed Kubernetes, Managed Slurm on Kubernetes, inference platform services, and AIOps capabilities.
  • Define networking and storage requirements for AI workloads, including CNI integration, high-performance fabrics, RDMA, GPUDirect, and managed infrastructure architecture.
  • Design self-healing systems, incident-response automation, chaos engineering programs, upgrade automation, security patching, and zero-downtime maintenance.
  • Set technical direction, lead design reviews, mentor engineers, influence cross-team infrastructure decisions, and represent Lambda through technical talks, blog posts, and customer engagements.

Requirements

  • 10+ years of experience in software engineering, platform engineering, or SRE, including at least 5 years focused on Kubernetes at scale.
  • Expert understanding of Kubernetes internals, including API machinery, controllers, schedulers, operators, CRDs, CSI, CNI, and extension patterns.
  • Holistic expertise across compute, networking, storage, and security, with experience designing distributed systems and managed multi-tenant platforms.
  • Strong production software engineering skills in Go and Python.
  • Deep experience with GPU orchestration in Kubernetes, including NVIDIA GPU Operator, device plugins, DCGM, MIG, time-slicing, and GPU-aware scheduling.
  • Proven technical leadership experience driving design decisions, mentoring engineers, and influencing infrastructure direction across teams.
  • Experience with observability at scale, including metrics, dashboards, distributed tracing, and actionable alerting.
  • Solid Linux systems and L2-L7 networking knowledge, including RDMA, InfiniBand, and RoCE.
  • Experience with infrastructure-as-code and GitOps workflows.
  • Preferred qualifications include experience with managed Kubernetes services or control-plane components; NVIDIA networking and GPU ecosystem projects; Slurm and Kubernetes-native batch schedulers; confidential computing; infrastructure migrations; CNCF, Kubernetes, or NVIDIA open-source contributions; multi-tenant security and compliance; and ML infrastructure.

Benefits

  • The role requires working from the San Francisco, San Jose, or Bellevue office four days per week, with Tuesday designated as the work-from-home day.
  • Generous cash and equity compensation is offered.
  • Health, dental, and vision coverage are provided for employees and dependents.
  • Wellness and commuter stipends are available for select roles.
  • A 401(k) plan with a 2% company match is available to U.S. employees.
  • Flexible paid time off is provided.

Tech Stack

GoGrafanaKubernetesLinuxPrometheusPython

Categories

Lambda

About Lambda

501-1,000 employees

Lambda provides GPU cloud computing and on-prem AI hardware—servers, clusters, and workstations—for teams training and serving large ML models. Its products include NVIDIA H100/A100 instances, managed clusters, and the Lambda Stack software, sold via usage-based cloud pricing and hardware sales. Founded in 2012 and headquartered in San Francisco, the privately held company serves researchers, startups, enterprises, and hyperscalers.

Contact me