3 months ago
Remote, WorldwideSenior
Responsibilities
- Partner with service teams to define customer-focused SLIs, SLOs, and error budget policies.
- Own and evolve the Operational Readiness Review process for new services and major changes.
- Connect postmortem findings to operational readiness gaps and drive systemic fixes for recurring failures.
- Provide reliability expertise for architecture reviews, failure-mode analysis, dependency mapping, and resilience design.
- Identify and quantify operational toil and build or advocate for automation to reduce it.
- Help teams develop sustainable on-call practices, including alert quality, escalation paths, runbook coverage, and noise reduction.
- Track organization-wide operational maturity, identify systemic gaps, and drive remediation.
Requirements
- At least 7 years of experience in SRE, production engineering, or reliability-focused roles, including shaping SRE practices and driving adoption across engineering teams.
- Strong software engineering mindset with the ability to write code and build tools.
- Hands-on experience defining and operationalizing SLOs and SLIs at scale, including error budget policies.
- Deep experience with incident response, postmortem facilitation, and systemic operational improvements.
- Experience with large-scale multi-tenant systems; managed database platforms or Postgres is a bonus.
- Proficiency with cloud infrastructure, preferably AWS, and infrastructure-as-code, preferably Pulumi; Terraform or CDK are also acceptable.
- Clear and persuasive communication skills for influencing without authority across a distributed organization.
- Experience working in async or globally distributed teams.
- Experience with Kubernetes-based platform operations is a plus.
- Familiarity with OpenTelemetry, VictoriaMetrics, Grafana, or similar observability tools is a plus.
- Experience building developer-facing reliability tooling such as SLO dashboards, ORR frameworks, toil tracking, or DORA metrics is a plus.
Benefits
- Fully remote with global hiring and a WeWork membership or coworking allowance.
- Employee ESOP equity ownership.
- Tech allowance for equipment and workspace setup.
- Health insurance covered at 100% for employees and 80% for dependents.
- Annual company off-site.
- Asynchronous and flexible work environment.
- Annual professional development and education allowance.
Tech Stack
Categories
DevOpsSite Reliability
About Supabase
Supabase builds an open-source Postgres-based backend platform for developers, bundling database, authentication, storage, realtime APIs, edge functions, and vector search. It monetizes through a managed cloud service and enterprise offerings while remaining deployable self-hosted. Founded in 2020 and privately held, the company operates globally with a fully remote team. Developers use Supabase as a Firebase alternative to launch quickly and scale production applications on PostgreSQL.
