Lambda

Senior Software Engineer - Core Cloud Platform

Lambda
Apply
2 months ago
San Francisco, CA, USA or San Jose, CA, USASenior
H1B sponsor

Base Salary

$296k - $346k/yr

Responsibilities

  • Build and operate core cloud platform services for compute lifecycle, bare-metal hosts, capacity, placement, and maintenance workflows.
  • Design APIs, backend services, state machines, and orchestration systems for Lambda’s GPU cloud.
  • Develop bare-metal lifecycle workflows for launch, termination, restart, host reclaim, validation, quarantine, and return to pool.
  • Improve deployment, observability, testing, alerting, runbooks, and operational readiness for business-critical services.
  • Debug production issues across distributed services, infrastructure dependencies, networking, and cloud workflows.
  • Partner with infrastructure, networking, fleet, security, support, and product teams on cross-system contracts and end-to-end capabilities.
  • Contribute to architecture, design documents, code reviews, incident follow-through, and mentoring.

Requirements

  • Bachelor’s degree or equivalent working experience.
  • 6+ years of professional software engineering experience building production backend or distributed systems.
  • Strong programming ability in Python, Go, or a similar backend or systems language.
  • Experience designing and operating APIs, workflow engines, schedulers, orchestration services, or other distributed systems.
  • Understanding of fault tolerance, idempotency, retries, state machines, failure handling, and production debugging.
  • Experience with cloud or cloud-like infrastructure primitives such as compute, networking, storage, capacity management, identity, or fleet operations.
  • Comfort with Linux, containers, Kubernetes, infrastructure automation, and service deployment patterns.
  • Experience owning production services, participating in on-call, and improving systems through operational learnings.
  • Experience with testability, observability, metrics, logging, alerting, and supportable operations.
  • Ability to drive ambiguous infrastructure problems through designs, implementation plans, and production outcomes.
  • Clear communication across engineering, product, support, infrastructure teams, and leadership.
  • Preferred experience with cloud control planes, compute platforms, schedulers, infrastructure orchestration, bare metal, GPU infrastructure, HPC, Kubernetes, Slurm, or large-scale AI/ML infrastructure.
  • Preferred experience with host lifecycle, provisioning, validation, firmware, BMC/Redfish, fleet management, or maintenance workflows.
  • Familiarity with Temporal, Airflow, event-driven systems, or durable workflow orchestration is preferred.
  • Preferred networking experience with VPCs, firewalls, SDN, InfiniBand, routing, or distributed-systems networking.
  • Security-minded engineering experience involving identity, authorization, audit logging, attestation, or tenant isolation is preferred.

Benefits

  • San Francisco/San Jose office presence 4 days per week, with Tuesday designated as the work-from-home day.
  • Health, dental, and vision coverage for employees and dependents.
  • Wellness and commuter stipends for select roles.
  • 401(k) plan.
  • Flexible paid time off plan.
  • Cash and equity compensation.

Tech Stack

Categories

Lambda

About Lambda

501-1,000 employees

Lambda provides GPU cloud computing and on-prem AI hardware—servers, clusters, and workstations—for teams training and serving large ML models. Its products include NVIDIA H100/A100 instances, managed clusters, and the Lambda Stack software, sold via usage-based cloud pricing and hardware sales. Founded in 2012 and headquartered in San Francisco, the privately held company serves researchers, startups, enterprises, and hyperscalers.

Contact me