4 days ago
Bengaluru, IndiaStaff+
Responsibilities
- Own platform reliability practices for availability, resilience, latency, and operational efficiency.
- Drive DevOps, automation, Golden Image, support automation, CI/CD, and GitHub Actions reliability initiatives.
- Lead JFrog Helm chart automation and JFrog image and Azure Container Registry migration work.
- Support microservices deployment enablement and platform and tooling upgrades.
- Own monitoring, alerting, observability, and logging components including Prometheus, AlertManager, Grafana, Azure Monitor, Thanos, OpenSearch, and FluentBit.
- Support health-check frameworks, including Airflow health-check requirements.
- Provide high-complexity incident troubleshooting support to Tier 1 and Tier 2 teams.
- Collaborate with architecture and delivery teams on reliability and scalability patterns.
- Lead cloud infrastructure creation, maintenance, governance, and access controls.
- Drive capacity planning, disaster recovery planning and exercises, platform documentation, cost management, role enforcement, and license governance.
- Lead high-severity incident response and execute post-incident reliability improvements.
- Maintain standard operating procedures for established alerts and incident patterns.
Requirements
- 6+ years of experience in SRE, platform engineering, DevOps, or advanced production support roles.
- Strong hands-on expertise with Kubernetes, particularly Azure Kubernetes Service, and cloud-native platform operations.
- Advanced CI/CD engineering and GitHub Actions experience.
- Deep observability experience with Prometheus, Grafana, AlertManager, and logging stacks.
- Strong Python automation scripting skills for reliability engineering, platform tooling, and operational toil reduction.
- End-user proficiency with AI-assisted productivity and operations tools; AI/ML model development is not required.
- Familiarity with Java, React, and Spring Boot services for production troubleshooting and stability improvements rather than feature development.
- Enterprise operational experience with Confluent Kafka, Confluent Cloud, Azure Event Hub, AWS-MSK, and Apache Flink.
- Experience with access management, role enforcement, separation of duties, and other governance controls.
- Proven high-severity incident leadership and post-incident reliability improvement execution.
- Postgres performance and reliability operations experience is preferred.
- Telecom-scale high-availability systems experience is preferred.
- The role is described as Senior to Lead IC, typically requiring 10 to 17 years of experience.
Benefits
- Opportunity to define and scale platform reliability standards.
- High technical ownership and strong cross-functional influence.
- Enterprise-scale impact across observability, automation, and resilience engineering.
- Regular full-time role with 40 weekly hours.
- Onsite work in Hyderabad, Bangalore, or a designated AT&T location.
Tech Stack
Apache AirflowApache FlinkGitHub ActionsGrafanaHelmJavaKubernetesPostgreSQLPrometheusPythonReactSpring Boot
Categories
DevOpsSite Reliability
