
SRE Architect (68019)
Hitachi, Ltd.2 months ago
London, United KingdomSenior
Responsibilities
- Define and enforce non-functional requirements for performance, scalability, availability, fault tolerance, and cost efficiency using FMEA-based failure analysis.
- Design and implement self-healing automation for known failure patterns to reduce human intervention and on-call burden.
- Build observability stacks with metrics, logs, traces, ML-driven anomaly detection, and AIOps capabilities.
- Automate incident detection, triage, escalation, remediation, and post-incident reviews.
- Automate database provisioning, release management, backup and restore, and operational request workflows.
- Define and track SLIs, SLOs, and error budgets for critical services.
- Conduct chaos engineering exercises and game days to validate resiliency.
- Mentor two SRE Engineers and establish engineering standards and a culture of reliability.
- Collaborate with Platform Engineering and Cloud teams to embed reliability into infrastructure and deployment pipelines.
Requirements
- 7+ years of experience in SRE, DevOps, or production engineering, including 3+ years in a senior or lead capacity.
- Proven experience improving availability, reducing MTTR, and implementing self-healing at scale.
- Experience managing or automating database operations in enterprise environments.
- Expertise with observability, resiliency, service management, reliability, performance engineering, scalability, release management, and cloud cost management.
- Strong SQL skills and experience with automated database provisioning, migration tools, and database release pipelines.
- Proficiency in Python, Go, and Bash automation and self-healing runbooks.
- Knowledge of Kubernetes, cloud platforms, networking, and storage systems.
- Experience with CI/CD and release engineering tools for integrated database and application releases.
- Experience with FinOps principles, resource optimization, and cloud spend analysis.
- Strong leadership, mentoring, stakeholder management, communication, analytical, problem-solving, and prioritization skills.
- Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent experience.
- Relevant certifications such as CKA, AWS DevOps Professional, Azure DevOps Expert, or SRE Foundation are preferred.
- Published SRE, observability, or chaos engineering work; service mesh and distributed tracing experience; and financial services or regulated-industry experience are desirable.
Benefits
- Industry-leading benefits, support, and services focused on holistic health and wellbeing.
- Flexible working arrangements depending on role and location.
- An inclusive, equal-opportunity workplace with reasonable accommodations available during recruitment.
- Autonomy, freedom, ownership, knowledge sharing, and a focus on work-life balance.
Tech Stack
AWSAzureBashDatadogFlywayGitLab CI/CDGoGoogle Cloud PlatformGrafanaIstioJenkinsKubernetesPrometheusPythonSpinnakerSQL
Categories
Site Reliability