Deloitte Advisory

Senior Team Lead | Site Reliability Engineering | Bengaluru | Engineering

Deloitte Advisory
Apply
24 hours ago
Bengaluru, IndiaSenior / Staff+

Responsibilities

  • Own end-to-end reliability, availability, scalability, cost, and performance for mission-critical production systems.
  • Improve MTTR, MTTA, incident response, SLO adherence, and operational efficiency through automation, runbooks, and process enhancements.
  • Participate in rotational 24x7 on-call coverage and manage high-severity incidents through resolution and documented follow-up.
  • Define and manage SLIs, SLOs, SLAs, error budgets, alerting strategies, and operational metrics.
  • Design, deploy, and manage AWS and GCP infrastructure, including Kubernetes environments, networking, IAM, load balancing, certificates, logging, and metrics.
  • Implement infrastructure as code with Terraform and manage Kubernetes workloads, deployments, storage, networking, Helm, YAML, Canary releases, and Blue-Green releases.
  • Build and maintain Jenkins pipelines and develop automation using Python, Shell, Groovy, and GitHub.
  • Operate observability and monitoring solutions using Dynatrace, Grafana, logs, metrics, and traces.
  • Troubleshoot distributed systems, containerized workloads, Java and Golang applications, infrastructure issues, and network-related problems.
  • Collaborate with engineering, operations, and cloud management teams to improve system design, resilience, and production debugging practices.
  • Identify and eliminate repetitive manual work while driving reliability engineering practices and culture.

Requirements

  • Bachelor's degree in Engineering.
  • 5-8 years of relevant and progressive experience in SRE, DevOps, or cloud engineering.
  • Hands-on experience managing production-grade systems in 24x7 environments.
  • Strong expertise in AWS, GCP, Kubernetes/GKE, networking, IAM, load balancing, certificates, KMS, logging, metrics exploration, BigQuery, and Pub/Sub.
  • Strong hands-on Terraform experience, including writing and debugging Terraform code from scratch.
  • Deep Kubernetes and Docker experience with strong troubleshooting capability in containerized environments.
  • Hands-on Jenkins and GitHub experience plus strong Python and Shell scripting skills.
  • Experience with Dynatrace, Grafana, and log-, metric-, and trace-based monitoring.
  • Working knowledge of Java and/or Golang applications and strong debugging skills across application and infrastructure layers.
  • Deep Linux and TCP/IP networking fundamentals, including the ability to debug distributed-system network issues.
  • Solid understanding of SLI, SLO, SLA, and error budgets, with demonstrable experience improving MTTR and MTTA.
  • Experience handling the incident management lifecycle and working effectively in high-pressure production environments.
  • Preferred experience includes high-scale distributed systems, banking or financial services, security and compliance practices, and Canary or Blue-Green deployment strategies.

Benefits

  • Bengaluru/Bangalore-based engineering role supporting 24x7 production operations with rotational on-call coverage.
  • The posting does not state specific compensation, benefits, or remote/hybrid work arrangements.

Tech Stack

Categories

DevOpsSite Reliability
Deloitte Advisory

About Deloitte Advisory

10,000+ employees

Deloitte Advisory provides consulting and managed services that help organizations manage risk, strengthen controls, investigate fraud, meet regulatory requirements, and execute transactions using analytics and industry expertise. Offerings span cybersecurity, M&A due diligence, finance and regulatory transformation, internal audit, and tech-enabled platforms. It is part of Deloitte’s global network (founded in 1845 and privately held), with U.S. member-firm teams serving corporations and government agencies across sectors.

Contact me