
Principal Site Reliability Engineer
commercetools15 days ago
London, United KingdomStaff+
Responsibilities
- Standardize end-to-end incident management processes, including detection, response, communication, and postmortems.
- Develop real-time metrics, dashboards, and signals to monitor system health and incident trends.
- Analyze operational data to identify process gaps and partner with product engineering teams on improvements.
- Lead organization-wide readiness programs for peak-traffic events such as Black Friday.
- Identify resiliency gaps, lead cross-team initiatives, and deliver technical outcomes with infrastructure and product teams.
- Drive communication, documentation, training, and knowledge sharing around resiliency and operational excellence.
Requirements
- 7+ years of experience driving incident management or operational excellence.
- 5+ years of experience leading organization-wide resiliency and reliability initiatives.
- Proven experience managing and scaling systems during high-stakes, multi-team operational events such as major launches or Black Friday.
- Strong ability to analyze metrics, diagnose technical issues, and measure process improvements.
- Experience evaluating both technical issues and organizational or human dynamics.
- Demonstrated success managing large-scale initiatives across multiple engineering teams in an Agile environment.
- Fluent English and exceptional written and verbal communication skills.
- Experience running technical training or onboarding is preferred.
- Strong self-awareness, customer focus, mentoring ability, and willingness to learn new technologies.
Benefits
- Comprehensive health benefits for employees and dependents, including access to OpenUp mental health support.
- Annual learning budget, self-paced learning platforms, language training, personalized coaching, mentorship, and leadership programs.
- Family Leave Plus with additional fully paid parental leave weeks beyond government-provided leave.
- Equity participation program.
- Hybrid work arrangement with three days per week in a Berlin, London, Munich, or Valencia office.
Categories
Site Reliability