4 days ago
Base Salary
$230k - $346k/yr
Responsibilities
- Design, build, and maintain scalable Kubernetes control-plane services, operators, and custom controllers.
- Develop Go and Python automation for end-to-end cluster provisioning, upgrades, patching, and deletion.
- Build GPU-aware orchestration systems supporting GPU scheduling and resource allocation.
- Partner with networking teams on CNI integration, high-performance fabrics, RDMA, and GPUDirect for AI workloads.
- Develop resilient distributed systems with timeouts, retries, backoff, and degraded-mode operation.
- Build inference platform services, including model serving infrastructure, inference-based autoscaling, and multi-model deployment patterns.
- Create internal tools and command-line interfaces for ML and AI teams to deploy and monitor inference services.
- Support and debug production issues through an on-call rotation.
Requirements
- At least 6 years of software engineering experience with significant technical ownership, such as driving projects from design through production or serving as a de facto technical lead.
- Deep understanding of Kubernetes internals, including controllers, schedulers, operators, CRDs, CSI, CNI, and extension patterns.
- Strong understanding of distributed-systems fundamentals, fault tolerance, graceful degradation, and failure handling at scale.
- Experience operating control planes and low-level components of large-scale Kubernetes clusters.
- Experience with observability at scale, including Prometheus, Grafana, distributed tracing, and actionable alerting systems.
- Strong programming skills in Go and Python and ability to collaborate on shared codebases.
- Knowledge of Linux systems, networking, containers, and cloud infrastructure.
- Experience building or operating managed Kubernetes services such as GKE, EKS, or AKS, or working on Kubernetes control-plane components is preferred.
- Hands-on experience with NVIDIA GPU and networking technologies such as GPU Operator, device plugins, DCGM, MIG, Network Operator, NCCL tuning, or similar is preferred.
- Familiarity with Slurm, KAI, Volcano, or Kueue and GPU, InfiniBand, RDMA, high-performance computing, or AI/ML storage architecture is preferred.
- Contributions to CNCF projects or Kubernetes SIGs are a plus.
Benefits
- Health, dental, and vision coverage for employees and dependents.
- Wellness and commuter stipends for select roles.
- 401(k) plan with a 2% company match for USA employees.
- Flexible paid time off plan.
- Generous cash and equity compensation.
- Hybrid work arrangement requiring presence in the San Francisco, San Jose, or Bellevue office four days per week, with Tuesday designated as the work-from-home day.
About Lambda
Lambda provides GPU cloud computing and on-prem AI hardware—servers, clusters, and workstations—for teams training and serving large ML models. Its products include NVIDIA H100/A100 instances, managed clusters, and the Lambda Stack software, sold via usage-based cloud pricing and hardware sales. Founded in 2012 and headquartered in San Francisco, the privately held company serves researchers, startups, enterprises, and hyperscalers.
