
Site Reliability Engineer III
JPMorgan Chase21 hours ago
Mumbai, IndiaSenior
Responsibilities
- Design reliability solutions, guide peers, and promote site reliability engineering practices.
- Collaborate with software engineers and partner teams on deployment and reliability approaches using CI/CD pipelines.
- Implement infrastructure, configuration, and network as code and improve solutions through incremental changes.
- Configure, maintain, monitor, and optimize applications and infrastructure for availability, reliability, and scalability.
- Use SLIs and SLOs to identify and address issues before customer impact.
- Use and validate enterprise-authorized AI capabilities for incident triage, troubleshooting, and post-incident analysis while following data sensitivity requirements.
- Identify reliability risks and recurring toil, prioritize reuse-focused improvements, and explore appropriate technologies.
- Participate in incident response, root-cause analysis, postmortems, problem management, operational standards, and mentoring.
Requirements
- Formal training or certification in site reliability engineering concepts and 3+ years of applied experience, subject to country-specific requirements.
- Proficiency in SRE principles and implementing SRE within an application or platform.
- Proficiency in at least one programming language such as Python, Java/Spring Boot, or .NET, with experience developing, debugging, and maintaining code in a large corporate environment.
- Experience validating AI-assisted operational recommendations and handling sensitive operational data appropriately.
- Experience with white-box and black-box monitoring, SLO alerting, telemetry collection, and troubleshooting common networking technologies and issues.
- Experience with continuous integration and continuous delivery tooling, containers, and container orchestration.
- Preferred: production support, DevOps, platform, on-call, incident response, RCA, postmortem, problem management, toil reduction, mentoring, runbooks, alerting, and SLO definition experience.
- Preferred: Dynatrace, Grafana, Prometheus, Splunk, OpenTelemetry, distributed tracing, alert-as-code, synthetic monitoring, and telemetry-based capacity planning experience.
- Preferred: hands-on Kubernetes and AWS operations, including EKS, IAM, and VPC, plus Terraform, Jenkins, GitLab CI, Helm, Kustomize, Istio, Linkerd, OPA, Gatekeeper, Vault, ArgoCD, and Flux.
- Preferred: Java microservices, Python or shell scripting, Linux/Unix troubleshooting, Kafka or IBM MQ, Oracle or MongoDB, certificate management, and microservice reliability patterns.
Tech Stack
Apache KafkaAWSGitLab CI/CDGrafanaHelmIstioJavaJenkinsKubernetesLinuxMongoDB.NETPrometheusPythonSplunkSpring BootTerraformVault
Categories
Site Reliability
About JPMorgan Chase
JPMorgan Chase provides consumer and commercial banking, payments, credit card, wealth management, and corporate and investment banking services to individuals, businesses, institutions, and governments. The public company (NYSE: JPM) earns revenue from interest, fees, trading, and asset management across operations in more than 100 markets. Headquartered in New York City with roots dating to 1799, it serves retail customers and prominent corporate and government clients through brands including Chase and J.P. Morgan.