
Sr. Site Reliability Engineer
MeridianLink2 months ago
Base Salary
$104k - $178k/yr
Responsibilities
- Design and maintain SLOs and SLIs for critical systems and ensure reliability targets are met.
- Lead monitoring, logging, tracing, and broader observability strategy.
- Build runbooks, incident response procedures, and post-incident review processes while mentoring the team on incident management.
- Architect and deploy highly available, recoverable cloud infrastructure on AWS or Azure using infrastructure as code.
- Develop automation and AIOps capabilities for incident detection, intelligent alerting, toil reduction, and self-healing systems.
- Drive reliability improvements through load testing, chaos engineering, failure analysis, and elimination of single points of failure.
- Partner with application and backend teams on reliable system design, architecture reviews, and reliability assessments.
- Write production-grade Python tooling for automation, metrics collection, alert management, and operational workflows.
- Champion infrastructure security, compliance, and defense-in-depth practices in a regulated fintech environment.
Requirements
- 7+ years of experience in Site Reliability Engineering, DevOps, platform engineering, or closely related production-system roles.
- Expert-level experience with AWS or Azure and infrastructure at scale, including compute, networking, storage, and managed services.
- Hands-on expertise designing monitoring, alerting, logging, and distributed tracing solutions with observability platforms.
- Strong knowledge of SLOs, SLIs, SLAs, error budgets, distributed systems, and highly available and scalable architectures.
- Proficiency in Python, PowerShell, bash, or similar scripting languages for production automation and operational tooling.
- Experience with AIOps practices such as event correlation, intelligent alerting, predictive analytics, and automated remediation.
- Experience with infrastructure-as-code tools, version control, and CI/CD pipeline design.
- Track record of incident management and on-call ownership, plus the ability to mentor engineers and influence cross-functionally.
- Preferred: experience in fintech, payments, banking, or other regulated industries and familiarity with SOC 2 or PCI-DSS compliance.
- Preferred: Kubernetes and container orchestration experience.
- Preferred: observability as code, custom metrics and dashboards, chaos engineering, reliability testing, or tools such as Gremlin.
- Preferred: open-source observability or infrastructure contributions, security expertise, and database optimization and backup/recovery experience.
Tech Stack
Categories
DevOpsSite Reliability
About MeridianLink
MeridianLink builds cloud software for financial institutions to open accounts, originate consumer and mortgage loans, manage collections, and verify data. Its platform is sold to banks, credit unions, and consumer reporting agencies via subscriptions with implementation and managed services. Founded in 1998 and headquartered in Irvine, CA, MeridianLink is a public company on the NYSE.