
Sr Lead SRE - Reliability Engineering & Problem Management
JPMorgan Chase12 hours ago
Hyderābād, IndiaStaff+
Responsibilities
- Serve as the technical authority for Problem Management-led root-cause analysis reviews of major incidents and service-impacting events.
- Lead deep-dive RCA challenge sessions and validate technical findings, causal chains, assumptions, and corrective actions through evidence-based analysis.
- Evaluate detection gaps, monitoring effectiveness, observability, automation opportunities, resilience weaknesses, process breakdowns, and human factors.
- Review vendor, third-party, and internal investigation reports and validate corrective-action closure and effectiveness.
- Define and govern Service Level Objectives, Service Level Indicators, reliability metrics, and error budgets.
- Evaluate system architecture against reliability principles and recommend resilience patterns such as graceful degradation, dependency isolation, rate limiting, circuit breakers, fault tolerance, capacity management, auto-remediation, and self-healing.
- Lead reliability maturity assessments across platforms and services.
- Drive enterprise-authorized AI workflows for RCA generation, incident analysis, log analytics, pattern discovery, problem trend analysis, and corrective-action recommendations.
- Establish governance for explainability, auditability, data handling, validation, resiliency, and security of AI-generated findings.
- Influence engineering stakeholders and translate technical findings into executive-ready narratives.
Requirements
- Formal training or certification in security engineering concepts and 5+ years of applied experience in Infrastructure Engineering, Site Reliability Engineering, Production Engineering, Systems Engineering, or Software Engineering.
- Experience leading or supporting critical incident investigations and conducting deep technical RCAs in enterprise-scale, highly regulated, and mission-critical environments.
- Advanced expertise in network engineering, cloud infrastructure, Linux and Windows platforms, middleware, databases, storage, application architecture, DevOps toolchains, distributed systems, and enterprise monitoring platforms.
- Hands-on expertise in SLO/SLI engineering, distributed tracing, telemetry design, reliability metrics, error-budget management, and AIOps platforms.
- Experience with Splunk, Dynatrace, Grafana, Datadog, Prometheus, AppDynamics, Elastic, and OpenTelemetry.
- Deep understanding of root-cause analysis methodologies, Five Whys, Fault Tree Analysis, event correlation, human factors analysis, systemic cause analysis, Problem Management Governance, and Major Incident Management.
- Demonstrated experience using enterprise-authorized AI capabilities to improve reliability workflows, with strong validation habits and awareness of data sensitivity.
- Ability to establish safe AI usage practices while maintaining resiliency, security, and auditability outcomes.
- Strong executive communication skills and the ability to constructively challenge senior engineering stakeholders and influence without direct authority.
- Preferred experience leading reliability engineering initiatives in large-scale complex environments, AI-enabled incident analysis, and governance for explainability and auditability of AI-generated findings.
Tech Stack
Categories
Site Reliability
About JPMorgan Chase
JPMorgan Chase provides consumer and commercial banking, payments, credit card, wealth management, and corporate and investment banking services to individuals, businesses, institutions, and governments. The public company (NYSE: JPM) earns revenue from interest, fees, trading, and asset management across operations in more than 100 markets. Headquartered in New York City with roots dating to 1799, it serves retail customers and prominent corporate and government clients through brands including Chase and J.P. Morgan.