
Principal Site Reliability Engineer
commercetools15 days ago
Valencia, SpainStaff+
Responsibilities
- Standardize company-wide incident detection, response, communication, and postmortem processes.
- Develop real-time metrics, dashboards, and signals for system health and incident trends.
- Analyze operational data to identify process gaps and partner with product engineering teams on improvements.
- Own and lead organization-wide readiness for major traffic spikes and high-stakes operational events.
- Identify resiliency gaps, collaborate with infrastructure and product teams, and deliver technical outcomes.
- Drive communication, documentation, training, and knowledge sharing around resiliency and operational excellence.
- Partner with engineering leadership, Staff Engineers, and Principal Engineers across cloud, security, API, architecture, and performance domains.
Requirements
- 7+ years of experience driving incident management or operational excellence.
- 5+ years leading organization-wide resiliency and reliability initiatives.
- Proven experience managing and scaling systems for high-stakes, multi-team operational events such as Black Friday or major launches.
- Strong ability to analyze metrics, diagnose technical issues, and measure process improvements.
- Ability to assess both technical defects and organizational or human dynamics.
- Demonstrated success managing large-scale initiatives spanning multiple engineering teams in an Agile environment.
- Fluent English with exceptional written and verbal communication skills.
- Experience running technical training or onboarding is preferred.
- Strong customer focus, self-awareness, mentoring ability, and willingness to learn new technologies.
Benefits
- Comprehensive health benefits for employees and dependents, including access to OpenUp mental health support.
- Annual learning budget, self-paced learning platforms, language training, personalized coaching, mentorship, and leadership programs.
- Additional fully paid parental leave through Family Leave Plus.
- Equity participation program.
- Hybrid work arrangement requiring three days per week in the Berlin, London, Munich, or Valencia office.
Categories
Site Reliability