BT

SRE Professional

BT
Apply
17 hours ago
Bengaluru, IndiaSenior

Responsibilities

  • Monitor service health, performance, availability, and operational metrics using observability platforms.
  • Investigate complex production incidents using application logs, database queries, transaction flows, and monitoring data.
  • Provide technical oversight for investigations, validate root causes, and ensure high-quality incident resolution.
  • Drive incident, problem, change, and release management activities, including root-cause analysis and preventive actions.
  • Partner with engineering, business, supplier, and support teams to deliver application, operational, and customer experience improvements.
  • Develop and maintain service dashboards, operational reports, and performance insights.
  • Identify and implement automation, self-service, AIOps, AI-driven operations, and self-healing capabilities.
  • Support resilience engineering, capacity planning, trend analysis, demand forecasting, and performance management for critical production services.

Requirements

  • Experience supporting or operating enterprise applications in complex production environments.
  • Strong understanding of application architecture, system integrations, data flows, and customer-facing business processes.
  • Ability to analyze application, middleware, and infrastructure logs for issue diagnosis and root-cause investigations.
  • Knowledge of SQL and Oracle databases, including production troubleshooting and performance tuning.
  • Hands-on experience with monitoring and observability platforms such as Dynatrace, including dashboard creation, alert analysis, and trend identification.
  • Understanding of Java-based applications, APIs, integration services, and AWS-hosted platforms.
  • Experience with incident, problem, change, and release management processes, including ServiceNow service operations.
  • Strong analytical and stakeholder management skills, including the ability to challenge technical investigations and influence engineering teams.
  • Experience identifying automation, resilience, and customer experience improvements through operational insights.
  • Exposure to Shell scripting or Python automation.
  • Experience using AI, ML, or GenAI capabilities for service reliability, observability, incident response, root-cause analysis, operational efficiency, and self-healing automation.
  • Knowledge of performance management, capacity planning, trend analysis, demand forecasting, and resilience engineering.
  • Experience with SRE principles, service reliability, availability, and operational excellence is desirable.

Categories

Site Reliability
BT

About BT

10,000+ employees

BT Group builds and operates telecom networks and digital services for consumers, enterprises, and public‑sector clients, selling broadband, fixed-line voice, mobile (via EE), and managed network, cloud, and cyber security services. It is headquartered in London and listed on the London Stock Exchange, with American depositary shares on the NYSE. BT also owns Openreach, which manages the UK’s fibre and copper access network used by multiple retail providers.

Contact me