15 days ago
Base Salary
$230k - $346k/yr
Responsibilities
- Design, build, and maintain Kubernetes control-plane services, operators, custom controllers, and end-to-end cluster lifecycle automation.
- Develop GPU-aware orchestration systems for scheduling and resource allocation.
- Integrate networking solutions for AI workloads, including CNI, high-performance fabrics, RDMA, and GPUDirect.
- Build resilient distributed systems with timeouts, retries, backoff, and degraded-mode operation.
- Develop inference platform services, model-serving infrastructure, inference-load autoscaling, and multi-model deployment patterns.
- Build internal tools and CLIs for ML and AI teams to deploy and monitor inference services.
- Support and debug production issues through an on-call rotation.
Requirements
- 6+ years of software engineering experience with ownership of significant technical scope.
- Deep understanding of Kubernetes internals, including controllers, schedulers, operators, CRDs, CSI, CNI, and extension patterns.
- Strong understanding of distributed-systems fundamentals, fault tolerance, graceful degradation, and failure handling.
- Experience operating control planes and low-level components of large-scale Kubernetes clusters.
- Experience with observability at scale, including Prometheus, Grafana, distributed tracing, and actionable alerting.
- Strong programming skills in Go and Python.
- Knowledge of Linux systems, networking, containers, and cloud infrastructure.
- Preferred experience building or operating managed Kubernetes services such as GKE, EKS, or AKS.
- Preferred hands-on experience with NVIDIA GPU and networking tools, including GPU Operator, device plugins, DCGM, MIG, Network Operator, or NCCL tuning.
- Familiarity with Slurm, KAI, Volcano, or Kueue and with GPU, InfiniBand, RDMA, or high-performance computing on Kubernetes.
- Exposure to storage architecture for AI/ML workloads and contributions to CNCF projects or Kubernetes SIGs is a plus.
Benefits
- Hybrid schedule requiring presence in the San Francisco, San Jose, or Bellevue office 4 days per week, with Tuesday designated as the work-from-home day.
- Generous cash and equity compensation.
- Health, dental, and vision coverage for employees and dependents.
- Wellness and commuter stipends for select roles.
- 401(k) plan with a 2% company match for USA employees.
- Flexible paid time off plan.
Tech Stack
Categories
About Lambda
Lambda provides GPU cloud computing and on-prem AI hardware—servers, clusters, and workstations—for teams training and serving large ML models. Its products include NVIDIA H100/A100 instances, managed clusters, and the Lambda Stack software, sold via usage-based cloud pricing and hardware sales. Founded in 2012 and headquartered in San Francisco, the privately held company serves researchers, startups, enterprises, and hyperscalers.
