5 days ago
Manila, PhilippinesStaff+
Responsibilities
- Define reliability engineering standards and the SLO framework.
- Prioritize systemic reliability risks and improvement work.
- Guide service-health analysis, capacity planning, and resilience practices.
- Partner with service owners while preserving remediation accountability.
- Own reliability engineering methods, SLO strategy, recurring-risk reduction, and improvement outcomes.
Requirements
- 10+ years of experience in SRE, reliability, production engineering, performance, or application diagnostics.
- Experience leading reliability programs or engineering practices.
- Strong service-impact, dependency, incident, and risk reasoning.
- Preferred experience in regulated or high-availability environments.
- Preferred experience with error budgets, capacity, resilience, or chaos engineering.
- Relevant SRE, cloud, or architecture certification is preferred.
Benefits
- Offers are contingent upon successfully passing a credit check, criminal background check, and pre-employment medical and drug screening.
Categories
Site Reliability
