2 months ago
Bellevue, WA, USA +2 moreMid Level
H1B sponsor
Base Salary
$266k - $395k/yr
Responsibilities
- Design, develop, and maintain performant, scalable, and reliable software for GPU and CPU compute infrastructure.
- Implement services for bare-metal and virtual-machine instancing.
- Develop distributed systems that manage and orchestrate compute resources across multiple SKUs.
- Build workflows for compute instance lifecycle, system health, host validation, and cluster validation.
- Contribute to internal platform tooling and improve bare-metal and VM developer and customer experiences.
- Troubleshoot and debug complex issues in production and development environments.
- Participate in on-call rotations and own incidents.
- Collaborate across engineering teams and help resolve ambiguous requirements and solutions through RFCs.
Requirements
- At least 3 years of production experience with Go or Python.
- At least 3 years of experience with bare-metal and virtualization hardware management and configuration.
- Comfort working in Linux environments and debugging at the operating system, hardware, and networking layers.
- Ability to independently troubleshoot complex systems and communicate across software, infrastructure, and vendor teams.
- Familiarity with GPU infrastructure or high-performance computing environments is preferred.
- Experience with Slurm or Kubernetes-based cluster management is preferred.
- Experience with public cloud internals, including virtualization, KVM, QEMU, security, and fleet health, is preferred.
- Experience with durable execution platforms such as Temporal is preferred.
Benefits
- Hybrid work arrangement requiring presence in the San Francisco, San Jose, or Bellevue office 4 days per week, with Tuesday designated as the work-from-home day
- Generous cash and equity compensation
- Health, dental, and vision coverage for employees and dependents
- Wellness and commuter stipends for select roles
- 401(k) plan with a 2% company match for U.S. employees
- Flexible paid time off plan
About Lambda
Lambda provides GPU cloud computing and on-prem AI hardware—servers, clusters, and workstations—for teams training and serving large ML models. Its products include NVIDIA H100/A100 instances, managed clusters, and the Lambda Stack software, sold via usage-based cloud pricing and hardware sales. Founded in 2012 and headquartered in San Francisco, the privately held company serves researchers, startups, enterprises, and hyperscalers.
