4 days ago
Base Salary
$314k - $465k/yr
Responsibilities
- Drive the technical vision and development of Lambda’s bare-metal Managed Kubernetes platform, including control-plane scalability, multi-tenancy, cluster lifecycle management, and high availability.
- Integrate and extend NVIDIA’s GPU orchestration ecosystem and design GPU-aware scheduling and placement systems.
- Lead development of managed orchestration services, including Managed Kubernetes, Managed Slurm on Kubernetes, inference platform services, and AIOps capabilities.
- Define networking and storage requirements for AI workloads, including CNI integration, high-performance fabrics, RDMA, GPUDirect, and managed infrastructure architecture.
- Design self-healing systems, incident-response automation, chaos engineering programs, upgrade automation, security patching, and zero-downtime maintenance.
- Set technical direction, lead design reviews, mentor engineers, influence cross-team infrastructure decisions, and represent Lambda through technical talks, blog posts, and customer engagements.
Requirements
- 10+ years of experience in software engineering, platform engineering, or SRE, including at least 5 years focused on Kubernetes at scale.
- Expert understanding of Kubernetes internals, including API machinery, controllers, schedulers, operators, CRDs, CSI, CNI, and extension patterns.
- Holistic expertise across compute, networking, storage, and security, with experience designing distributed systems and managed multi-tenant platforms.
- Strong production software engineering skills in Go and Python.
- Deep experience with GPU orchestration in Kubernetes, including NVIDIA GPU Operator, device plugins, DCGM, MIG, time-slicing, and GPU-aware scheduling.
- Proven technical leadership experience driving design decisions, mentoring engineers, and influencing infrastructure direction across teams.
- Experience with observability at scale, including metrics, dashboards, distributed tracing, and actionable alerting.
- Solid Linux systems and L2-L7 networking knowledge, including RDMA, InfiniBand, and RoCE.
- Experience with infrastructure-as-code and GitOps workflows.
- Preferred qualifications include experience with managed Kubernetes services or control-plane components; NVIDIA networking and GPU ecosystem projects; Slurm and Kubernetes-native batch schedulers; confidential computing; infrastructure migrations; CNCF, Kubernetes, or NVIDIA open-source contributions; multi-tenant security and compliance; and ML infrastructure.
Benefits
- The role requires working from the San Francisco, San Jose, or Bellevue office four days per week, with Tuesday designated as the work-from-home day.
- Generous cash and equity compensation is offered.
- Health, dental, and vision coverage are provided for employees and dependents.
- Wellness and commuter stipends are available for select roles.
- A 401(k) plan with a 2% company match is available to U.S. employees.
- Flexible paid time off is provided.
Tech Stack
Categories
About Lambda
Lambda provides GPU cloud computing and on-prem AI hardware—servers, clusters, and workstations—for teams training and serving large ML models. Its products include NVIDIA H100/A100 instances, managed clusters, and the Lambda Stack software, sold via usage-based cloud pricing and hardware sales. Founded in 2012 and headquartered in San Francisco, the privately held company serves researchers, startups, enterprises, and hyperscalers.
