2 months ago
Base Salary
$296k - $346k/yr
Responsibilities
- Build and operate core cloud platform services for compute lifecycle, bare-metal hosts, capacity, placement, and maintenance workflows.
- Design APIs, backend services, state machines, and orchestration systems for Lambda’s GPU cloud.
- Develop bare-metal lifecycle workflows for launch, termination, restart, host reclaim, validation, quarantine, and return to pool.
- Improve deployment, observability, testing, alerting, runbooks, and operational readiness for business-critical services.
- Debug production issues across distributed services, infrastructure dependencies, networking, and cloud workflows.
- Partner with infrastructure, networking, fleet, security, support, and product teams on cross-system contracts and end-to-end capabilities.
- Contribute to architecture, design documents, code reviews, incident follow-through, and mentoring.
Requirements
- Bachelor’s degree or equivalent working experience.
- 6+ years of professional software engineering experience building production backend or distributed systems.
- Strong programming ability in Python, Go, or a similar backend or systems language.
- Experience designing and operating APIs, workflow engines, schedulers, orchestration services, or other distributed systems.
- Understanding of fault tolerance, idempotency, retries, state machines, failure handling, and production debugging.
- Experience with cloud or cloud-like infrastructure primitives such as compute, networking, storage, capacity management, identity, or fleet operations.
- Comfort with Linux, containers, Kubernetes, infrastructure automation, and service deployment patterns.
- Experience owning production services, participating in on-call, and improving systems through operational learnings.
- Experience with testability, observability, metrics, logging, alerting, and supportable operations.
- Ability to drive ambiguous infrastructure problems through designs, implementation plans, and production outcomes.
- Clear communication across engineering, product, support, infrastructure teams, and leadership.
- Preferred experience with cloud control planes, compute platforms, schedulers, infrastructure orchestration, bare metal, GPU infrastructure, HPC, Kubernetes, Slurm, or large-scale AI/ML infrastructure.
- Preferred experience with host lifecycle, provisioning, validation, firmware, BMC/Redfish, fleet management, or maintenance workflows.
- Familiarity with Temporal, Airflow, event-driven systems, or durable workflow orchestration is preferred.
- Preferred networking experience with VPCs, firewalls, SDN, InfiniBand, routing, or distributed-systems networking.
- Security-minded engineering experience involving identity, authorization, audit logging, attestation, or tenant isolation is preferred.
Benefits
- San Francisco/San Jose office presence 4 days per week, with Tuesday designated as the work-from-home day.
- Health, dental, and vision coverage for employees and dependents.
- Wellness and commuter stipends for select roles.
- 401(k) plan.
- Flexible paid time off plan.
- Cash and equity compensation.
About Lambda
Lambda provides GPU cloud computing and on-prem AI hardware—servers, clusters, and workstations—for teams training and serving large ML models. Its products include NVIDIA H100/A100 instances, managed clusters, and the Lambda Stack software, sold via usage-based cloud pricing and hardware sales. Founded in 2012 and headquartered in San Francisco, the privately held company serves researchers, startups, enterprises, and hyperscalers.
