about 4 hours ago
Remote, WorldwideSenior
Responsibilities
- Own the reliability of Supabase's deployment and release systems against clear SLOs and error budgets.
- Standardize and instrument fragmented deployment workflows to create a trustworthy pre-production signal.
- Drive disaster-recovery readiness by making environments reproducibly deployable from scratch.
- Build and operate health and SLO monitoring for critical user flows using synthetic testing.
- Reduce mean-time-to-detect and mean-time-to-recover for deploy-related incidents.
- Participate in on-call duties and lead blameless postmortems to improve operational practices.
- Document operational procedures to ensure reliability knowledge is accessible.
- Define and track SLAs, SLOs, error budgets, and DORA delivery metrics.
Requirements
- 5+ years of experience in SRE, production operations, platform engineering, or release engineering.
- Experience operating production systems at scale and carrying on-call responsibilities.
- Fluency in SLAs, SLOs, error budgets, DORA metrics, and observability tooling.
- Experience leading incident response and driving down MTTD/MTTR.
- Proficiency in AWS, infrastructure-as-code, and Kubernetes.
- Strong scripting and automation skills to eliminate toil.
- Excellent communication skills with both infrastructure specialists and product engineers.
- Ability to thrive in async, globally distributed teams and navigate ambiguity.
Benefits
- Fully remote work with a WeWork membership or co-working allowance.
- Equity ownership through ESOP for all team members.
- Tech allowance to set up an ideal work environment.
- 100% health insurance coverage for employees and 80% for dependents.
- Annual off-sites for team connection and collaboration.
- Flexible work hours with an emphasis on asynchronous operations.
- Annual education allowance for professional development.
