
Vice President - Lead Site Reliability Engineer
JPMorgan Chase16 days ago
Responsibilities
- Champion site reliability culture, practices, knowledge sharing, and reuse-first reliability workflows.
- Lead initiatives that improve application and platform reliability, stability, performance, scalability, and service levels.
- Define service level indicators, service level objectives, and error budgets with stakeholders.
- Serve as the primary technical contact during major incidents and minimize business impact through rapid troubleshooting and resolution.
- Build and govern AI-assisted operational workflows for triage, runbook execution, incident summarization, change validation, capacity analysis, and post-incident review.
- Collect and analyze monitoring and telemetry data and implement dashboards across test and production environments.
- Identify and remediate capacity risks, platform interdependencies, and technology bottlenecks with application and infrastructure teams.
- Provide technical leadership, mentorship, detailed incident write-ups, and guidance across multiple technical domains.
Requirements
- Formal training or certification in site reliability engineering concepts and 5+ years of applied experience.
- Hands-on proficiency in reliability, scalability, performance, security, enterprise architecture, toil reduction, and SRE practices.
- Fluency in at least one programming language and experience applying AI-assisted automation and agentic patterns to engineering or operational workflows.
- Proficiency in scripting, automation, infrastructure-as-code, and cloud technologies across public and private environments.
- Experience using enterprise-authorized AI for incident investigation, knowledge capture, capacity analysis, and other SRE workflows, with strong validation and data-sensitivity practices.
- Ability to evaluate AI-assisted recommendations, establish operational guardrails, and align outcomes with resiliency and security expectations.
- Advanced observability, monitoring, alerting, and telemetry experience, including production dashboard implementation.
- Experience with continuous integration and delivery, containers, container orchestration, networking, Linux or Windows operating systems, databases, and deployment practices.
- Preferred qualifications include AWS, Azure, or GCP experience or certifications; GitHub-based code reviews; Terraform; GitHub Copilot or similar coding assistants; AI agents for infrastructure operations; and support for complex mission-critical applications.
- Familiarity with modern front-end technologies is preferred.
Benefits
- Competitive total rewards package with base salary determined by role, experience, skills, and location, plus eligible commission or discretionary incentive compensation.
- Comprehensive health care coverage, on-site health and wellness centers, retirement savings plan, backup childcare, tuition reimbursement, mental health support, and financial coaching.
- Equal opportunity employer with reasonable accommodation support and a stated commitment to diversity and inclusion.
Tech Stack
Categories
Site Reliability
About JPMorgan Chase
JPMorgan Chase provides consumer and commercial banking, payments, credit card, wealth management, and corporate and investment banking services to individuals, businesses, institutions, and governments. The public company (NYSE: JPM) earns revenue from interest, fees, trading, and asset management across operations in more than 100 markets. Headquartered in New York City with roots dating to 1799, it serves retail customers and prominent corporate and government clients through brands including Chase and J.P. Morgan.