
Site Reliability Engineer II
Early Warning Services2 days ago
Chicago, IL, USA +2 moreMid Level
Base Salary
$83k - $132k/yr
Responsibilities
- Improve the reliability, resilience, scalability, observability, recoverability, performance, and operational health of production services.
- Define and implement SLIs, SLOs, error budgets, monitoring, alerting, dashboards, logging, tracing, and service-health instrumentation.
- Drive improvements across deployment practices, Infrastructure as Code, automation, testing, incident response, capacity management, resilience, and operational readiness.
- Identify systemic production issues and improve code, architecture, automation, tooling, and engineering practices.
- Partner with Software Engineering teams to integrate reliability and operational readiness throughout the development lifecycle.
- Participate in or lead incident response, troubleshooting, service restoration, and post-incident learning.
- Participate independently in an on-call rotation and diagnose and resolve production incidents and service degradation.
- Reduce operational toil through software, automation, reusable patterns, and improved engineering practices.
Requirements
- Typically 2–5 years of relevant professional experience in Software Engineering, Site Reliability Engineering, Systems Engineering, Cloud/Platform Engineering, DevOps, Infrastructure Engineering, or a comparable technical discipline.
- Experience with software development or scripting using one or more modern programming languages.
- Experience with software engineering principles, distributed systems, production troubleshooting, automation, and observability.
- Experience with public cloud technologies and architectures, preferably AWS, as well as infrastructure, networking, Linux/Unix, and modern application architectures.
- Experience developing, deploying, operating, or improving highly available production software or distributed systems.
- Experience with Infrastructure as Code, containers or orchestration, observability, monitoring, alerting, software-delivery automation, incident management, performance analysis, capacity management, resilience testing, disaster recovery, or operational readiness.
- Experience creating reusable automation, tooling, platforms, patterns, or practices that improve engineering effectiveness.
- Bachelor's degree in Computer Science, Software Engineering, Computer Engineering, Information Systems, or a related technical field, or equivalent practical experience.
- Demonstrated analytical, problem-solving, communication, and collaboration skills.
Benefits
- Hybrid work model.
- Base pay ranges from $83,000–$110,000 in Phoenix or Chicago and $99,000–$132,000 in San Francisco, plus eligibility for a discretionary incentive plan.
- Medical, dental, and vision coverage, with Health Savings Account and flexible spending account options.
- 401(k) plan with a 100% company Safe Harbor Match on the first 6% of deferral immediately upon eligibility.
- Flexible Time Off for exempt employees or generous PTO for non-exempt employees, 11 paid company holidays, and a paid volunteer day.
- 12 weeks of paid parental leave.
- Maven Family Planning support for fertility, adoption, surrogacy, pregnancy, postpartum, pediatrics, and returning to work.
Tech Stack
Categories
Site Reliability
About Early Warning Services
Early Warning Services builds payment networks and fraud/risk mitigation solutions for U.S. banks and credit unions. It operates the Zelle peer-to-peer payments network and is rolling out Paze, a digital wallet, alongside identity verification and account risk services. Headquartered in Scottsdale, Arizona, it is owned by a consortium of major U.S. banks and partners with thousands of financial institutions nationwide.