Hitachi, Ltd.

SRE Architect (68019)

Hitachi, Ltd.
Apply
2 months ago
London, United KingdomSenior

Responsibilities

  • Define and enforce non-functional requirements for performance, scalability, availability, fault tolerance, and cost efficiency using FMEA-based failure analysis.
  • Design and implement self-healing automation for known failure patterns to reduce human intervention and on-call burden.
  • Build observability stacks with metrics, logs, traces, ML-driven anomaly detection, and AIOps capabilities.
  • Automate incident detection, triage, escalation, remediation, and post-incident reviews.
  • Automate database provisioning, release management, backup and restore, and operational request workflows.
  • Define and track SLIs, SLOs, and error budgets for critical services.
  • Conduct chaos engineering exercises and game days to validate resiliency.
  • Mentor two SRE Engineers and establish engineering standards and a culture of reliability.
  • Collaborate with Platform Engineering and Cloud teams to embed reliability into infrastructure and deployment pipelines.

Requirements

  • 7+ years of experience in SRE, DevOps, or production engineering, including 3+ years in a senior or lead capacity.
  • Proven experience improving availability, reducing MTTR, and implementing self-healing at scale.
  • Experience managing or automating database operations in enterprise environments.
  • Expertise with observability, resiliency, service management, reliability, performance engineering, scalability, release management, and cloud cost management.
  • Strong SQL skills and experience with automated database provisioning, migration tools, and database release pipelines.
  • Proficiency in Python, Go, and Bash automation and self-healing runbooks.
  • Knowledge of Kubernetes, cloud platforms, networking, and storage systems.
  • Experience with CI/CD and release engineering tools for integrated database and application releases.
  • Experience with FinOps principles, resource optimization, and cloud spend analysis.
  • Strong leadership, mentoring, stakeholder management, communication, analytical, problem-solving, and prioritization skills.
  • Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent experience.
  • Relevant certifications such as CKA, AWS DevOps Professional, Azure DevOps Expert, or SRE Foundation are preferred.
  • Published SRE, observability, or chaos engineering work; service mesh and distributed tracing experience; and financial services or regulated-industry experience are desirable.

Benefits

  • Industry-leading benefits, support, and services focused on holistic health and wellbeing.
  • Flexible working arrangements depending on role and location.
  • An inclusive, equal-opportunity workplace with reasonable accommodations available during recruitment.
  • Autonomy, freedom, ownership, knowledge sharing, and a focus on work-life balance.

Tech Stack

AWSAzureBashDatadogFlywayGitLab CI/CDGoGoogle Cloud PlatformGrafanaIstioJenkinsKubernetesPrometheusPythonSpinnakerSQL

Categories

Site Reliability
Hitachi, Ltd.

About Hitachi, Ltd.

10,000+ employees
Contact me