AT&T

Senior App/Prod Support (Tier 3 SRE for event-driven ecosystems)

AT&T
Apply
4 days ago
Bengaluru, IndiaStaff+

Responsibilities

  • Own platform reliability practices for availability, resilience, latency, and operational efficiency.
  • Drive DevOps, automation, Golden Image, support automation, CI/CD, and GitHub Actions reliability initiatives.
  • Lead JFrog Helm chart automation and JFrog image and Azure Container Registry migration work.
  • Support microservices deployment enablement and platform and tooling upgrades.
  • Own monitoring, alerting, observability, and logging components including Prometheus, AlertManager, Grafana, Azure Monitor, Thanos, OpenSearch, and FluentBit.
  • Support health-check frameworks, including Airflow health-check requirements.
  • Provide high-complexity incident troubleshooting support to Tier 1 and Tier 2 teams.
  • Collaborate with architecture and delivery teams on reliability and scalability patterns.
  • Lead cloud infrastructure creation, maintenance, governance, and access controls.
  • Drive capacity planning, disaster recovery planning and exercises, platform documentation, cost management, role enforcement, and license governance.
  • Lead high-severity incident response and execute post-incident reliability improvements.
  • Maintain standard operating procedures for established alerts and incident patterns.

Requirements

  • 6+ years of experience in SRE, platform engineering, DevOps, or advanced production support roles.
  • Strong hands-on expertise with Kubernetes, particularly Azure Kubernetes Service, and cloud-native platform operations.
  • Advanced CI/CD engineering and GitHub Actions experience.
  • Deep observability experience with Prometheus, Grafana, AlertManager, and logging stacks.
  • Strong Python automation scripting skills for reliability engineering, platform tooling, and operational toil reduction.
  • End-user proficiency with AI-assisted productivity and operations tools; AI/ML model development is not required.
  • Familiarity with Java, React, and Spring Boot services for production troubleshooting and stability improvements rather than feature development.
  • Enterprise operational experience with Confluent Kafka, Confluent Cloud, Azure Event Hub, AWS-MSK, and Apache Flink.
  • Experience with access management, role enforcement, separation of duties, and other governance controls.
  • Proven high-severity incident leadership and post-incident reliability improvement execution.
  • Postgres performance and reliability operations experience is preferred.
  • Telecom-scale high-availability systems experience is preferred.
  • The role is described as Senior to Lead IC, typically requiring 10 to 17 years of experience.

Benefits

  • Opportunity to define and scale platform reliability standards.
  • High technical ownership and strong cross-functional influence.
  • Enterprise-scale impact across observability, automation, and resilience engineering.
  • Regular full-time role with 40 weekly hours.
  • Onsite work in Hyderabad, Bangalore, or a designated AT&T location.

Tech Stack

Apache AirflowApache FlinkGitHub ActionsGrafanaHelmJavaKubernetesPostgreSQLPrometheusPythonReactSpring Boot

Categories

DevOpsSite Reliability
Contact me