
Principal Site Reliability Engineer
PowerPlan, Inc.11 days ago
Atlanta, GA, USAStaff+
Responsibilities
- Resolve escalated infrastructure cases across AWS and Azure.
- Build automation and tooling to eliminate repetitive operational work and reduce manual resolution time.
- Establish incident response and blameless post-incident review processes for critical production incidents.
- Build and operate an SLO-aligned observability platform with dashboards, tuned alerts, and reliability reporting.
- Coach teams on incident communication and decision-making.
- Influence engineering practices without formal authority.
Requirements
- Deep hands-on experience operating production systems in AWS and Azure environments.
- Strong operational automation skills using Python and PowerShell.
- Experience identifying and eliminating repetitive operational work through automation.
- Experience leading incident response and blameless post-incident reviews.
- Strong observability expertise, particularly with Grafana and SLI/SLO-driven monitoring.
- Extensive experience in cloud operations, site reliability engineering, or infrastructure engineering, or equivalent professional experience.
- Clear written and verbal communication skills with technical and non-technical audiences.
Benefits
- Hybrid work arrangement combining onsite work at the corporate office with work from home.
- Flexible working arrangements may be accommodated when sensible, with onsite attendance required for scheduled office days, team meetings, client meetings, or special events.