
Sr. Site Reliability Engineer - Digital Assets
Early Warning Services1 day ago
Chicago, IL, USA +2 moreSenior
Base Salary
$106k - $156k/yr
Responsibilities
- Improve the reliability, resilience, scalability, performance, recoverability, observability, and operational health of production services.
- Define and implement SLIs, SLOs, error budgets, dashboards, monitoring, alerting, logging, tracing, and other service-health instrumentation.
- Drive improvements across CI/CD, Infrastructure as Code, automation, testing, incident response, capacity management, resilience, and operational readiness.
- Identify systemic production issues and improve code, architecture, automation, tooling, and engineering practices.
- Partner with Software Engineering teams to incorporate reliability throughout the development lifecycle.
- Participate in or lead incident response, troubleshooting, service restoration, and blameless post-incident learning.
- Reduce operational toil through software, automation, reusable patterns, and improved engineering practices.
- Mentor engineers, share knowledge, develop reusable solutions, and guide less-experienced engineers across complex reliability problems.
Requirements
- Typically 5–8 years of relevant professional experience in Software Engineering, Site Reliability Engineering, Systems Engineering, Cloud/Platform Engineering, DevOps, Infrastructure Engineering, or a comparable technical discipline.
- Experience with software development or scripting using one or more modern programming languages.
- Experience with software engineering principles, distributed systems, production troubleshooting, automation, and observability.
- Experience with public cloud technologies and architectures, preferably AWS, along with infrastructure, networking, Linux/Unix, and modern application architectures.
- Experience developing, deploying, operating, or improving highly available production software or distributed systems.
- Experience with CI/CD, Infrastructure as Code, containers or orchestration, observability, monitoring, alerting, and software-delivery automation.
- Experience with SLIs, SLOs, error budgets, incident management, performance analysis, capacity management, resilience testing, disaster recovery, or operational readiness.
- Experience creating reusable automation, tooling, platforms, patterns, or practices that improve engineering effectiveness.
- Bachelor’s degree in Computer Science, Software Engineering, Computer Engineering, Information Systems, or a related technical field, or equivalent practical experience.
- Demonstrated analytical, problem-solving, communication, collaboration, software engineering, systems thinking, troubleshooting, and production reliability skills.
Benefits
- Hybrid work model for positions in Scottsdale, San Francisco, Chicago, or New York.
- Base pay ranges by location: $106,000–$130,000 in Phoenix/Chicago and $128,000–$156,000 in San Francisco, plus a discretionary incentive plan.
- Medical, dental, and vision coverage, with Health Savings Account and flexible spending account options.
- 401(k) plan with a 100% company Safe Harbor Match on the first 6% deferred, immediately upon eligibility.
- Flexible time off for exempt employees, generous PTO for non-exempt employees, 11 paid company holidays, and a paid volunteer day.
- 12 weeks of paid parental leave and Maven Family Planning support.
- Normal office environment with primarily sedentary computer-based work and occasional physical activities including lifting up to 10 pounds.
Tech Stack
Categories
DevOpsSite Reliability
About Early Warning Services
Early Warning Services builds payment networks and fraud/risk mitigation solutions for U.S. banks and credit unions. It operates the Zelle peer-to-peer payments network and is rolling out Paze, a digital wallet, alongside identity verification and account risk services. Headquartered in Scottsdale, Arizona, it is owned by a consortium of major U.S. banks and partners with thousands of financial institutions nationwide.