about 4 hours ago
Base Salary
$240k - $356k/yr
Responsibilities
- Operate and scale critical platform services across Lambda’s data centers.
- Improve the reliability of compute provisioning and orchestration systems.
- Build monitoring, alerting, and tracing for service health and failures.
- Define SLIs, SLOs, error budgets, and operational readiness standards.
- Automate detection and remediation of configuration drift and failed workflows.
- Build safe deployment, rollback, and disaster recovery workflows.
- Design fault-isolation mechanisms to prevent cascading failures.
- Lead production incident response and postmortems.
- Partner with various teams including Compute, Networking, and Security.
- Participate in on-call duties and improve sustainability through automation.
- Mentor engineers and raise the reliability bar across the organization.
Requirements
- 7+ years of experience in site reliability, infrastructure, or production software engineering.
- Deep experience operating Kubernetes in production environments.
- Understanding of Kubernetes architecture and common failure modes.
- Experience with physical data centers or hybrid cloud environments.
- Proficient with Terraform or similar infrastructure-as-code tools.
- Experience building CI/CD or GitOps workflows using relevant tools.
- Familiarity with observability platforms like Prometheus or Grafana.
- Ability to build production-quality tooling in Go, Python, or similar languages.
- Understanding of distributed systems concepts including consistency and retries.
- Experience defining and operating against SLIs and SLOs.
- Strong communication skills and ability to work effectively across teams.
Benefits
- Generous cash and equity compensation.
- Health, dental, and vision coverage for you and your dependents.
- Wellness and commuter stipends for select roles.
- 401k Plan with 2% company match for USA employees.
- Flexible paid time off plan that is actively encouraged.
