Lead Site Reliability Engineer, Electronic Colo Trading
JPMorgan ChaseResponsibilities
- Define non-functional requirements, availability targets, service-level indicators, service-level objectives, and error budgets for applications and platforms.
- Lead initiatives to improve reliability, stability, scalability, performance, and operational excellence through data-driven analysis and toil reduction.
- Optimize latency, throughput, and stability across Linux kernels, CPU isolation, NUMA, IRQs, NICs, networks, and trading infrastructure.
- Operate and improve metrics, logging, alerting, observability, runbooks, and service-level improvement processes.
- Lead major incident response, including triage, mitigation, root-cause analysis, corrective actions, and prevention of recurring incidents.
- Coordinate server break/fix work, remote hands, firmware alignment, maintenance windows, and standardized rack-level practices in colocation environments.
- Automate system administration and reliability workflows, including AI-assisted incident investigation, validation, knowledge capture, and operational readiness.
- Serve as the primary technical contact during major incidents and provide technical guidance and mentorship to other engineers.
Requirements
- Bachelor’s degree in Computer Science, Cybersecurity, Data Science, or a related discipline.
- Formal training or certification in site reliability engineering concepts and 5+ years of applied experience.
- Demonstrated proficiency in reliability, scalability, performance, security, enterprise architecture, toil reduction, and SRE practices.
- Professional production Linux administration experience in mission-critical environments, including RHEL-derived or Debian/Ubuntu systems.
- Strong troubleshooting skills across operating system, hardware, and basic networking layers, including Linux internals and performance tuning.
- Proficiency in infrastructure automation using Bash and Python or Go, plus experience with configuration management, monitoring, observability, and root-cause analysis.
- Experience using enterprise-authorized AI capabilities for SRE workflows, evaluating recommendations for correctness and risk, and applying data-sensitivity and security controls.
- Preferred experience supporting electronic trading or other latency-sensitive, high-availability environments; colocation data centers; low-latency Linux tuning; infrastructure-as-code and golden server builds; and large-scale fleet management, patching, and configuration drift control.
About JPMorgan Chase
With a history tracing its roots to 1799 in New York City, JPMorganChase is one of the world's oldest, largest, and best-known financial institutions—carrying forth the innovative spirit of our heritage firms in global operations across 100 markets. We serve millions of customers and many of the world’s most prominent corporate, institutional, and government clients daily, managing assets and investments, offering business advice and strategies, and providing innovative banking solutions and services. Social Media Terms and Conditions: https://bit.ly/JPMCSocialTerms JPMorgan Chase & Co. is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or status as a protected veteran.