JPMorgan Chase

Lead Site Reliability Engineer

JPMorgan Chase
Apply
11 hours ago
Bengaluru, IndiaStaff+

Responsibilities

  • Engineer reliability into enterprise-scale platforms through production software, automation, control loops, self-healing systems, and operational tooling.
  • Build declarative, intent-based systems that continuously reconcile desired state with reality.
  • Instrument and analyze telemetry and state data to support detection, diagnosis, and closed-loop remediation.
  • Define and operationalize SLIs, SLOs, error budgets, SLO-based alerting, telemetry standards, and actionable observability.
  • Own services end to end, including reliability, performance, security, cost, operability, and observability.
  • Participate in on-call rotations and lead triage, mitigation, communications, and blameless post-incident reviews during major incidents.
  • Reduce operational toil through measurable automation and durable engineering fixes.
  • Lead resiliency design reviews, break complex problems into executable work, serve as a technical lead for medium- to large-sized products, and mentor engineers.
  • Use and govern AI-assisted development, testing, incident analysis, and reliability workflows with appropriate validation, security, traceability, and auditability.
  • Establish reliability standards and guide partner organizations in secure, resilient, and operationally effective engineering practices.

Requirements

  • Formal training or certification in site reliability engineering concepts and at least 5 years of applied experience.
  • Strong production coding skills in an industry-standard language such as Python, Go, Java, C++, or Rust.
  • Hands-on experience operating production systems at scale, including on-call ownership, incident response, and reliability or operability design.
  • Practical experience with SLIs, SLOs, and error budgets, or clear aptitude to own and evolve them.
  • Deep observability experience including white-box and black-box monitoring, SLO-based alerting, and telemetry.
  • Proficiency with *nix systems and infrastructure automation and tooling such as Kubernetes and Terraform.
  • Strong systems thinking across interfaces, contracts, failure modes, and interactions at scale.
  • Ability to direct AI tools to perform substantive engineering work and exercise sound judgment about appropriate AI use.
  • Security-first, outcome-oriented judgment across reliability, impact, operational risk, and cost.
  • Experience using enterprise-authorized AI capabilities to improve SRE workflows, with strong validation practices and awareness of data sensitivity.
  • Ability to evaluate AI-assisted operational recommendations, define team guardrails, and align outcomes with resiliency and security expectations.
  • Preferred qualifications include networking depth, experience with network-adjacent platforms, multiple infrastructure domains or programming languages, ongoing AI skill development, AI-enabled workflow redesign, regulated or large-scale enterprise experience, and experience establishing engineering culture.

Tech Stack

C++DatadogGoGrafanaJavaKubernetesPrometheusPythonRustSplunkTerraform

Categories

Site Reliability
JPMorgan Chase

About JPMorgan Chase

10,000+ employees

JPMorgan Chase provides consumer and commercial banking, payments, credit card, wealth management, and corporate and investment banking services to individuals, businesses, institutions, and governments. The public company (NYSE: JPM) earns revenue from interest, fees, trading, and asset management across operations in more than 100 markets. Headquartered in New York City with roots dating to 1799, it serves retail customers and prominent corporate and government clients through brands including Chase and J.P. Morgan.

Contact me