
Senior Site Reliability Engineer
Inspire Brands22 days ago
Atlanta, GA, USASenior
Responsibilities
- Define and manage SLIs, SLOs, and error budgets for critical services.
- Drive production readiness reviews and incorporate reliability requirements into architecture and design.
- Perform capacity planning, failure mode analysis, dependency risk assessments, and systemic reliability remediation.
- Design monitoring, alerting, logging, tracing, dashboards, and telemetry to improve service observability.
- Lead technical responses to high-severity incidents and conduct blameless postmortems focused on systemic fixes.
- Improve incident detection, response, recovery, on-call processes, and alert signal-to-noise ratio.
- Automate repetitive operational work and build self-healing systems, tooling, and scripts.
- Improve deployment safety and support infrastructure as code.
- Conduct load testing, performance benchmarking, bottleneck analysis, and scalability improvements.
- Partner with engineering teams on resiliency patterns and mentor engineers on SRE best practices.
Requirements
- At least 5 years of experience in Site Reliability Engineering, Software Engineering, or Platform Engineering.
- At least 2 years of experience with Kubernetes and containerized workloads.
- A four-year degree in Computer Science or a related field.
- Strong programming or scripting skills in Python, Go, Java, or Node.js.
- Experience defining and operating against SLOs and error budgets.
- Experience leading incident response and root cause analysis for production systems.
- Strong understanding of distributed systems and microservices architecture.
- Deep expertise in at least one major cloud platform: Azure, AWS, or GCP.
- Expertise with observability platforms and monitoring strategy.
- Preferred: experience with chaos engineering, resiliency testing, high-volume transactional systems, AI-assisted observability or operational automation, and internal SRE tooling or platforms.
Categories
Site Reliability