
Sr. Site Reliability Engineer
MeridianLinkabout 3 hours ago
Base Salary
$104k - $178k/yr
Responsibilities
- Design, implement, and maintain Service Level Objectives (SLOs) and Service Level Indicators (SLIs).
- Lead observability strategy by designing monitoring, logging, and tracing architectures.
- Build and own runbooks, incident response procedures, and post-incident review processes.
- Architect and deploy cloud infrastructure on AWS or Azure with high availability.
- Develop automation and AIOps capabilities to reduce toil and enable self-healing systems.
- Drive reliability improvements through load testing and chaos engineering.
- Partner with application teams to design reliable systems from inception.
- Write production-grade Python tooling for automation and operational workflows.
- Champion security and compliance in infrastructure.
Requirements
- 7+ years in Site Reliability Engineering, DevOps, or related roles.
- Expert-level experience with Azure or AWS and managing infrastructure at scale.
- Demonstrated expertise in observability and hands-on experience with observability platforms.
- Strong background in SLOs, SLIs, and SLAs with experience defining meaningful objectives.
- Proven experience designing and troubleshooting highly available and scalable systems.
- Proficiency in Python, PowerShell, bash, and other scripting languages.
- Hands-on experience with AIOps practices and familiarity with AIOps platforms.
- Experience with infrastructure-as-code tools like Terraform or CloudFormation.
- Track record of incident management and on-call ownership.
- Excellent communication skills and ability to mentor junior engineers.