7 months ago
Responsibilities
- Analyze systemic failure patterns and design reliability improvements that prevent recurring incidents.
- Define and maintain SLO and SLA frameworks and use error budgets to guide reliability investments.
- Build automation, tooling, dashboards, and standards that reduce incident-response toil.
- Own Rootly configuration, workflows, and integrations with PagerDuty, Jira, Confluence, and Slack.
- Own incident management standards, serve as an on-call and escalation Incident Commander, and improve response practices.
- Train engineering teams, coach post-mortems, and guide teams in developing actionable corrective actions.
- Edit and review customer-facing incident documents and drive accurate, timely root cause analyses.
- Partner with engineering leaders as a trusted advisor on reliability practices and organizational change.
Requirements
- 10+ years of experience in SRE, incident management, or reliability engineering.
- Cloud experience with at least one of AWS, GCP, or Azure.
- Deep expertise with incident management tooling such as Rootly, PagerDuty, or similar platforms.
- Strong understanding of distributed systems and failure modes at scale; Kafka or event-streaming expertise is preferred.
- Deep observability experience spanning metrics, logging, and tracing.
- Experience with Kubernetes, container orchestration, CI/CD pipelines, and release processes.
- Familiarity with SLO and SLA frameworks, post-mortem facilitation, and reliability program development.
- Track record of advising engineering organizations and driving organization-wide process and cultural changes.
- Strong written communication skills for design documents, one-pagers, and runbooks.
- Experience collaborating asynchronously across time zones and navigating reliability programs in organizations with 500+ engineers.
- Multi-cloud experience across at least two of AWS, GCP, and Azure is an advantage.
- Experience with modern CI/CD, GitHub, and AI-assisted workflows is an advantage.
Benefits
- Global follow-the-sun team coverage with clean handoffs designed to support sustainable working hours.
- Inclusive, equal-opportunity workplace with opportunities for employees to lead, grow, and contribute across diverse perspectives.
Categories
Site Reliability
About Confluent
Confluent is pioneering a fundamentally new category of data infrastructure focused on data in motion. Our cloud-native offering is the foundational platform for data in motion --- designed to be the intelligent connective tissue enabling real-time data, from multiple sources, to constantly stream across the organization. With Confluent, our customers can meet the new business imperative of delivering rich, digital customer experiences and real-time business operations. Our mission is to help every organization harness data in motion so they can compete and thrive in the modern world.
