
Sr. Staff Site Reliability Engineer
Early Warning Services2 days ago
Scottsdale, AZ, USAStaff+
Base Salary
$150k - $200k/yr
Responsibilities
- Improve the reliability, resilience, scalability, performance, and operational health of production services.
- Define and improve SLIs, SLOs, error budgets, observability, monitoring, alerting, dashboards, and service-health instrumentation.
- Drive improvements across deployment practices, automation, Infrastructure as Code, testing, incident response, capacity management, resilience, and operational readiness.
- Identify systemic production issues and translate operational experience into improvements in code, architecture, tooling, and engineering practices.
- Partner with Software Engineering teams to incorporate reliability, observability, recoverability, scalability, and operational readiness throughout the development lifecycle.
- Lead or participate in incident response, troubleshooting, service restoration, post-incident learning, and sustainable on-call improvements.
- Provide broad technical leadership, escalation support, reusable solutions, and technical direction across a major domain or portfolio.
- Develop Staff and Senior engineers and create sustainable technical leadership depth.
Requirements
- Typically 12+ years of relevant professional experience in Software Engineering, Site Reliability Engineering, Systems Engineering, Cloud/Platform Engineering, DevOps, Infrastructure Engineering, or a comparable technical discipline.
- Experience with software development or scripting using one or more modern programming languages.
- Experience with software engineering principles, distributed systems, production troubleshooting, automation, and observability.
- Experience with public cloud technologies, infrastructure, networking, Linux/Unix, and modern application architectures.
- Demonstrated analytical, problem-solving, communication, and collaboration skills appropriate to the role's scope.
- Hands-on AWS experience is preferred, or comparable experience with Microsoft Azure, Google Cloud Platform, or Oracle Cloud Infrastructure.
- Experience developing, deploying, operating, or improving highly available production software or distributed systems.
- Experience with deployment automation, Infrastructure as Code, containers or orchestration, observability, monitoring, alerting, software-delivery automation, incident management, resilience testing, disaster recovery, or operational readiness.
- Bachelor's degree in Computer Science, Software Engineering, Computer Engineering, Information Systems, or a related technical field, or equivalent practical experience.
Benefits
- Hybrid work model for positions in Scottsdale, San Francisco, Chicago, or New York.
- Medical, dental, and vision coverage, plus HSA or FSA options.
- 401(k) plan with a 100% company Safe Harbor Match on the first 6% of deferral immediately upon eligibility.
- Flexible time off for exempt employees or generous PTO for non-exempt employees, 11 paid company holidays, and a paid volunteer day.
- 12 weeks of paid parental leave and Maven Family Planning support.
- Discretionary incentive plan and additional employee benefits.
Tech Stack
Categories
Site Reliability
About Early Warning Services
Early Warning Services builds payment networks and fraud/risk mitigation solutions for U.S. banks and credit unions. It operates the Zelle peer-to-peer payments network and is rolling out Paze, a digital wallet, alongside identity verification and account risk services. Headquartered in Scottsdale, Arizona, it is owned by a consortium of major U.S. banks and partners with thousands of financial institutions nationwide.