22 hours ago
Bengaluru, IndiaStaff+
Responsibilities
- Lead reliability engineering practices across applications, platforms, and infrastructure.
- Define and monitor SLIs, SLOs, error budgets, uptime targets, and operational KPIs.
- Own production stability, performance, resilience, service availability, and operational maturity.
- Lead production readiness reviews, reliability assessments, capacity planning, resilience engineering, and proactive service improvement.
- Provide expert third-line technical support and lead major incident resolution for critical business applications and platforms.
- Conduct blameless post-incident reviews, root cause analysis, and corrective and preventive actions.
- Design and maintain CI/CD pipelines, automated testing and validation, release controls, configuration management, and infrastructure-as-code practices.
- Implement enterprise observability solutions, dashboards, alerting, synthetic monitoring, and performance analytics.
- Contribute to reliability roadmaps, architectural improvements, platform engineering practices, and operational maturity initiatives.
- Mentor team members on DevOps and SRE practices.
- Collaborate with UI designers and frontend teams on accessible design systems, dashboard standards, component libraries, interaction guidelines, and usability improvements.
Requirements
- 10+ years of experience across SRE, DevOps, platform engineering, application support, production support, product engineering, UX design, enterprise application design, or operational tooling.
- Experience managing mission-critical enterprise applications and services in complex production environments.
- Strong understanding of SRE principles, including SLIs, SLOs, error budgets, observability, security, production readiness, and operational support processes.
- Hands-on experience with AWS, Azure, or GCP.
- Strong experience with Linux/Unix systems, networking, DNS, load balancers, certificates, firewalls, API gateways, middleware, APIs, databases, and cloud-native services.
- Experience with Kubernetes, Docker, CI/CD platforms, infrastructure as code, configuration management, and automation tooling.
- Experience with monitoring and observability tools such as Dynatrace, Prometheus, Grafana, ELK/EFK, Splunk, Datadog, AppDynamics, or New Relic.
- Strong scripting or programming experience with Python, Shell, PowerShell, Java, JavaScript, React, Angular, CSS, or similar technologies.
- Experience with production troubleshooting, root cause analysis, incident, problem, change, release management, and service improvement practices.
- Experience designing or improving enterprise applications, operational dashboards, admin portals, support tools, workflow systems, or data-heavy platforms.
- Proficiency with UX and design tools such as Figma, Axure, Sketch, or Adobe XD.
- Strong communication and stakeholder management skills; telecom domain knowledge is desirable.
- A bachelor's degree in Computer Science, Engineering, Information Technology, or a related discipline is desirable.
Benefits
- The posting lists the location as Bengaluru.
- Salary is described as competitive, with no explicit base-pay amount provided.
Tech Stack
AngularAWSAzureCSSDatadogDockerGoogle Cloud PlatformGrafanaJavaJavaScriptKubernetesLinuxPowerShellPrometheusPythonReactSplunk
Categories
DevOpsSite Reliability
About BT
BT Group builds and operates telecom networks and digital services for consumers, enterprises, and public‑sector clients, selling broadband, fixed-line voice, mobile (via EE), and managed network, cloud, and cyber security services. It is headquartered in London and listed on the London Stock Exchange, with American depositary shares on the NYSE. BT also owns Openreach, which manages the UK’s fibre and copper access network used by multiple retail providers.
