2 months ago
Base Salary
$206k - $275k/yr
Responsibilities
- Design, write, and deliver software and services that improve the availability, scalability, reliability, and efficiency of internal IT systems and platforms
- Build automation for mission-critical services to prevent recurring problems and automate responses to non-exceptional events
- Influence and create designs, architectures, standards, and methods for large-scale distributed systems
- Perform service capacity planning, demand forecasting, software performance analysis, and system tuning
- Produce documentation and related artifacts for owned systems
- Implement and manage paging, alerting, and on-call scheduling flows
Requirements
- Interest in system design and architecting for performance, scalability, and multiple cloud infrastructure platforms
- Knowledge of configuration management systems and toolchains such as Chef, Ansible, Terraform, and GitHub Actions
- Solid programming skills in Python, Go, or similar languages
- Ability to collaborate asynchronously and document issues and solutions
- Strong problem-solving, communication, and ownership mindset
Benefits
- Hybrid work arrangement requiring presence in the San Francisco or San Jose office 4 days per week, with Tuesday designated as the work-from-home day
- Generous cash and equity compensation
- Health, dental, and vision coverage for employees and dependents
- Wellness and commuter stipends for select roles
- 401(k) plan with a 2% company match for U.S. employees
- Flexible paid time off plan
Categories
DevOpsSite Reliability
About Lambda
Lambda provides GPU cloud computing and on-prem AI hardware—servers, clusters, and workstations—for teams training and serving large ML models. Its products include NVIDIA H100/A100 instances, managed clusters, and the Lambda Stack software, sold via usage-based cloud pricing and hardware sales. Founded in 2012 and headquartered in San Francisco, the privately held company serves researchers, startups, enterprises, and hyperscalers.
