4 days ago
Base Salary
$230k - $346k/yr
Responsibilities
- Design, build, and operate backend services, APIs, control planes, and platform capabilities for Lambda’s AI cloud.
- Solve distributed-systems challenges involving state, consistency, concurrency, scheduling, failure recovery, and lifecycle management.
- Own systems from problem framing and architecture through implementation, testing, rollout, observability, on-call, and continuous improvement.
- Improve availability, latency, throughput, efficiency, security, and operability at scale.
- Turn incidents and near misses into durable improvements through automation, testing, guardrails, and backstops.
- Collaborate with product, infrastructure, networking, storage, security, and SRE teams to resolve dependencies and deliver customer outcomes.
- Use AI-assisted development tools while independently verifying correctness, security, and maintainability.
- Contribute to technical standards, design and code reviews, mentorship, and—at Staff level—cross-team architecture.
Requirements
- Seven or more years of professional software engineering experience, or equivalent evidence of impact building production systems.
- Depth in at least one general-purpose programming language, with primary team usage of Go and Python.
- Experience designing, building, and operating backend services, distributed systems, infrastructure, or platform capabilities at meaningful scale.
- Practical understanding of system design, data models, APIs, failure modes, performance, and production reliability tradeoffs.
- Track record of owning complex work through delivery and operation, including testing, staged rollout, monitoring, incident response, and root-cause improvement.
- Proven ability to align cross-functional partners and gain consensus around decisions and tradeoffs.
- Preferred experience includes cloud services or platform infrastructure on AWS, GCP, Azure, or comparable platforms; Kubernetes and control-plane systems; infrastructure automation; durable workflows; event-driven architectures; infrastructure as code; highly available multi-region systems; and GPU, HPC, or AI/ML workloads.
Benefits
- Hybrid arrangement requiring presence in a San Francisco, San Jose, or Bellevue office four days per week, with Tuesday designated as the work-from-home day.
- Generous cash and equity compensation.
- Health, dental, and vision coverage for employees and dependents.
- Wellness and commuter stipends for select roles.
- 401(k) plan with a 2% company match for U.S. employees.
- Flexible paid time off plan.
About Lambda
Lambda provides GPU cloud computing and on-prem AI hardware—servers, clusters, and workstations—for teams training and serving large ML models. Its products include NVIDIA H100/A100 instances, managed clusters, and the Lambda Stack software, sold via usage-based cloud pricing and hardware sales. Founded in 2012 and headquartered in San Francisco, the privately held company serves researchers, startups, enterprises, and hyperscalers.
