28 days ago
Base Salary
$314k - $465k/yr
Responsibilities
- Drive the technical vision and development of Lambda’s bare-metal Managed Kubernetes platform, including control plane scalability, multi-tenancy, cluster lifecycle management, and high availability.
- Integrate and extend NVIDIA’s GPU orchestration and networking ecosystem, including GPU Operator, Network Operator, DCGM, NCCL, AICR, and Topograph.
- Design GPU-aware orchestration, scheduling, inference platform services, model serving infrastructure, autoscaling, and multi-model deployment patterns.
- Build the foundation for Managed Slurm on Kubernetes and support HPC workloads alongside Kubernetes workloads.
- Define networking and storage requirements for AI workloads, including CNI integration, high-performance fabrics, RDMA, and GPUDirect.
- Design self-healing systems, incident-response automation, root-cause analysis workflows, platform resilience, and chaos engineering programs.
- Establish operational excellence through upgrade automation, security patching, and zero-downtime maintenance.
- Set technical direction, lead design reviews, influence infrastructure roadmaps, mentor engineers, and standardize practices across teams.
- Partner with Network, Storage, Security, Customer Success, NVIDIA, customers, and the open-source community.
- Shape Lambda’s AIOps vision for capacity planning, anomaly detection, and predictive infrastructure maintenance.
- Represent Lambda through technical writing, conference talks, and strategic customer engagements.
Requirements
- 10+ years of experience in software engineering, platform engineering, or SRE, including at least 5 years focused on Kubernetes at scale.
- Expert understanding of Kubernetes internals, including API machinery, controllers, schedulers, operators, CRDs, CSI, CNI, and extension patterns.
- Holistic infrastructure expertise spanning compute, networking, storage, and security.
- Strong production software engineering skills in Go and Python.
- Deep experience with GPU orchestration in Kubernetes, including NVIDIA GPU Operator, device plugins, DCGM, MIG, time-slicing, and GPU-aware scheduling.
- Proven technical leadership experience driving cross-team decisions, mentoring engineers, and influencing infrastructure direction.
- Deep experience designing and operating managed services or multi-tenant platforms for external customers.
- Strong understanding of distributed systems, including consensus, fault tolerance, consistency models, and graceful degradation.
- Experience with observability at scale, including Prometheus, Grafana, distributed tracing, and actionable alerting.
- Knowledge of Linux systems, L2-L7 networking, RDMA, InfiniBand, and RoCE.
- Experience with infrastructure-as-code and GitOps workflows.
- Preferred experience building or operating managed Kubernetes services such as GKE, EKS, or AKS, or working on Kubernetes control plane components.
- Preferred experience with NVIDIA Network Operator, NCCL tuning, Topograph, AICR, or similar projects.
- Familiarity with Slurm, KAI, Volcano, or Kueue and traditional or Kubernetes-native batch scheduling.
- Background in confidential computing, ML infrastructure, customer migrations, or security and compliance in multi-tenant environments.
- Familiarity with RBAC, Pod Security Standards, network policies, workload isolation, CNCF projects, Kubernetes SIGs, or NVIDIA open-source projects.
Benefits
- Requires working from the San Francisco, San Jose, or Bellevue office 4 days per week, with Tuesday designated as the work-from-home day.
- Health, dental, and vision coverage for employees and dependents.
- Wellness and commuter stipends for select roles.
- 401(k) plan with company match for U.S. employees.
- Flexible paid time off plan.
- Cash and equity compensation are offered.
