2 months ago
Remote, WorldwideStaff+
Responsibilities
- Embed with development teams from design and architecture reviews through launch readiness to ensure reliability is designed into projects.
- Establish and evolve design reviews, launch checklists, and operational acceptance criteria and drive adoption across the organization.
- Define SLIs, SLOs, and error budgets with product and engineering teams and support data-driven prioritization.
- Create frameworks, tooling, and golden paths across AWS, GCP, and Azure using Terraform.
- Set technical direction and lead cross-team reliability initiatives as an individual contributor without direct people management.
- Partner with engineering, product, and design to align reliability work with business objectives and influence roadmaps.
- Mentor senior and mid-level engineers and improve operational excellence across teams.
- Lead incident management and blameless postmortems and turn operational signals into lasting improvements.
- Build relationships with cloud and infrastructure providers, influence their roadmaps, and support customer trust conversations.
- Participate in the shared SRE on-call rotation, currently one week per month.
Requirements
- Significant experience as an SRE, production engineer, platform engineer, or software engineer with a deep reliability focus, including experience operating at scale.
- Fluency across AWS, GCP, and Azure, with Terraform as the common infrastructure-as-code layer.
- Experience supporting or contributing to TypeScript frontend and backend services and Ruby on Rails backend services.
- Strong software engineering fundamentals and the ability to write production-quality code.
- Track record of changing how teams work and leading across team boundaries without formal authority.
- Ability to drive ambiguous, high-scope problems to completion with minimal oversight.
- Ability to identify organizational process, communication, and technical debt and propose improvements.
- Experience building measurement and evaluation frameworks, analyzing operational data, and translating findings into organizational improvements.
- Strong verbal and written English communication skills.
- A college degree is not required.
- Bonus: experience standing up or maturing an SRE practice at a growth-stage company.
- Bonus: experience as an embedded SRE or closely partnering with product teams.
- Bonus: experience designing chaos or resilience testing or progressive delivery practices.
Benefits
- Fully remote, globally distributed team
- Remote-friendly internationally; U.S. location is not required
- One week of on-call rotation per month
Tech Stack
Categories
DevOpsSite Reliability
