
Senior Site Reliability Engineer
Lloyds Banking Group16 hours ago
Hyderābād, IndiaSenior / Staff+
Responsibilities
- Design and continuously improve monitoring, alerting, logging, and observability solutions for reliable production services.
- Define and improve Service Level Indicators, Service Level Objectives, and Error Budgets.
- Investigate complex production incidents and support service restoration with Production Support and engineering teams.
- Support Post Incident Reviews and Problem Management by identifying reliability improvements and preventative engineering actions.
- Reduce operational toil through automation, tooling, and engineering improvements.
- Develop operational runbooks, support processes, and service-readiness practices.
- Embed reliability, resilience, monitoring, alerting, diagnostics, and operational documentation into the software development lifecycle.
- Identify reliability risks and deliver engineering improvements that reduce incidents and improve service resilience.
- Support onboarding of applications into the SRE operating model.
- Coach and mentor engineers, contribute to SRE standards and tooling, and provide line management where appropriate.
Requirements
- 9–15 years of experience.
- Strong understanding of Site Reliability Engineering principles and practices.
- Experience designing and implementing observability solutions, including monitoring, logging, and alerting.
- Experience defining and improving SLIs, SLOs, and Error Budgets.
- Strong troubleshooting and technical investigation skills across complex production environments.
- Experience with incident management, Post Incident Reviews, Problem Management, operational toil reduction, and automation.
- Experience using Infrastructure as Code and CI/CD tooling.
- Strong scripting or programming skills in one or more of Python, Java, JavaScript, PowerShell, or Bash.
- Experience with cloud technologies and modern application platforms.
- Strong understanding of cloud security, networking, and operational resilience.
- Strong stakeholder management, communication, collaboration, and cross-functional working skills.
- Experience with Kubernetes and containerised platforms is desirable.
- Experience with Azure, Google Cloud Platform, or AWS is desirable.
- Experience with Dynatrace, Grafana, Prometheus, ELK, or Splunk is desirable.
- Experience in Financial Services or another regulated industry is desirable.
- Experience with Jira, Confluence, and Agile delivery practices is desirable.
- Experience mentoring or coaching engineers and previous line management experience are desirable.
- Relevant Cloud, DevOps, or SRE certifications are desirable.
Benefits
- Hybrid working and flexible working options.
- Full-time role based in Hyderabad.
Tech Stack
Categories
DevOpsSite Reliability
About Lloyds Banking Group
Lloyds Banking Group is a UK-focused retail, commercial, and insurance banking group that provides current accounts, savings, mortgages, loans, credit cards, business banking, wealth, and pensions (via Scottish Widows). It earns revenue from interest margins and fees across digital and branch channels. Formed in 2009 and headquartered in London, it operates well-known brands including Lloyds Bank, Halifax, Bank of Scotland, and Scottish Widows, and is listed on the London Stock Exchange.