3 months ago
Base Salary
$160k - $220k/yr
Responsibilities
- Own the stability, robustness, performance, and production reliability of the platform across the full stack.
- Scale the on-call culture through stronger runbooks, escalation paths, and incident response processes.
- Develop monitoring, alerting, and observability capabilities so issues are detected and diagnosed quickly.
- Define and prioritize the SRE roadmap, including SLOs and chaos engineering investments.
- Reduce engineering toil and improve developer tooling, feedback loops, and deployment workflows.
- Collaborate with backend, ML, and frontend engineers to embed reliability best practices across squads.
Requirements
- Demonstrated senior-level experience in SRE, infrastructure, or backend engineering in production-critical environments.
- Hands-on experience with cloud infrastructure, preferably GCP; AWS or Azure experience is also considered.
- Experience with structured on-call processes, alerting policies, runbooks, incident postmortems, and escalation flows.
- Strong software engineering fundamentals and experience writing maintainable infrastructure-as-code, including Terraform.
- Ability to work autonomously, prioritize work, and drive initiatives to completion.
- Clear communication skills with engineers and leadership across distributed teams, including during incidents.
- Experience with Kubernetes, Cloud Run, GKE, Pub/Sub, PostgreSQL at scale, or healthcare security/compliance is a bonus.
Benefits
- Competitive salary and stock options.
- 100% individual coverage for medical, dental, and vision insurance.
- Unlimited paid time off, 11 national holidays, and unlimited sick leave.
- Paid parental leave.
- Remote-friendly work arrangement with $1,500 for home office equipment.
- Ownership of time and schedule.
- Regular exercise activities, off-sites, and team gatherings.
Categories
BackendSite Reliability