
Lead Site Reliability Engineer
JPMorgan Chase1 day ago
Responsibilities
- Build production-grade reliability software, including automation, control loops, self-healing systems, and tooling that eliminates manual operations.
- Design declarative, intent-based systems that continuously reconcile desired and actual state.
- Instrument services, collect telemetry at scale, and use state data for detection, diagnosis, and closed-loop remediation.
- Define and operationalize SLIs, SLOs, error budgets, SLO-based alerting, and actionable observability practices.
- Own services end to end across reliability, performance, security, and cost; lead on-call rotations, major-incident response, mitigation, communications, and blameless post-incident reviews.
- Lead resiliency design reviews, break complex problems into work for engineers, serve as technical lead for medium- to large-sized products, and mentor other engineers.
- Drive measurable toil reduction and reuse-first adoption of validated AI-assisted reliability workflows across the SDLC and toolchain.
- Evaluate AI-assisted operational recommendations, establish usage guardrails, and ensure resiliency, security, traceability, and auditability.
Requirements
- Formal training or certification in site reliability engineering concepts and 5+ years of applied experience.
- Experience using enterprise-authorized AI capabilities to improve SRE workflows, including incident investigation support and knowledge capture, with strong validation habits and data-sensitivity awareness.
- Ability to assess AI-assisted operational recommendations for correctness and risk and define appropriate team guardrails.
- Strong production software engineering experience in at least one of Python, Go, Java, C++, or Rust.
- Experience owning production systems at scale, including on-call responsibility, incident response, postmortems, and reliability or operability design.
- Experience with SLI/SLO and error-budget practices or demonstrated ability and motivation to institutionalize them.
- Deep observability experience with white-box and black-box monitoring, SLO-based alerting, and high-quality telemetry.
- Strong Unix fundamentals and infrastructure automation experience, including technologies such as Kubernetes, Terraform, and CI/CD.
- Experience with systems thinking at scale, service interfaces and contracts, failure modes, safe rollouts, and complex service dependencies.
- Security-first judgment, effective communication during high-severity events, and practical use of AI tools for engineering work.
- Preferred: networking expertise, experience across multiple infrastructure domains or programming languages, ongoing AI skill development, regulated or large-scale enterprise experience, and experience establishing engineering culture.
Benefits
- Competitive total rewards package with base salary determined by role, experience, skill set, and location; eligible roles may also receive commission-based or discretionary incentive compensation.
- Comprehensive health care coverage, on-site health and wellness centers, retirement savings plan, backup childcare, tuition reimbursement, mental health support, and financial coaching.
- Additional compensation and benefits details are provided during the hiring process.
Categories
DevOpsSite Reliability
About JPMorgan Chase
JPMorgan Chase provides consumer and commercial banking, payments, credit card, wealth management, and corporate and investment banking services to individuals, businesses, institutions, and governments. The public company (NYSE: JPM) earns revenue from interest, fees, trading, and asset management across operations in more than 100 markets. Headquartered in New York City with roots dating to 1799, it serves retail customers and prominent corporate and government clients through brands including Chase and J.P. Morgan.