PowerPlan, Inc.

Principal Site Reliability Engineer

PowerPlan, Inc.
Apply
11 days ago
Atlanta, GA, USAStaff+

Responsibilities

  • Resolve escalated infrastructure cases across AWS and Azure.
  • Build automation and tooling to eliminate repetitive operational work and reduce manual resolution time.
  • Establish incident response and blameless post-incident review processes for critical production incidents.
  • Build and operate an SLO-aligned observability platform with dashboards, tuned alerts, and reliability reporting.
  • Coach teams on incident communication and decision-making.
  • Influence engineering practices without formal authority.

Requirements

  • Deep hands-on experience operating production systems in AWS and Azure environments.
  • Strong operational automation skills using Python and PowerShell.
  • Experience identifying and eliminating repetitive operational work through automation.
  • Experience leading incident response and blameless post-incident reviews.
  • Strong observability expertise, particularly with Grafana and SLI/SLO-driven monitoring.
  • Extensive experience in cloud operations, site reliability engineering, or infrastructure engineering, or equivalent professional experience.
  • Clear written and verbal communication skills with technical and non-technical audiences.

Benefits

  • Hybrid work arrangement combining onsite work at the corporate office with work from home.
  • Flexible working arrangements may be accommodated when sensible, with onsite attendance required for scheduled office days, team meetings, client meetings, or special events.

Tech Stack

AWSAzureGrafanaPowerShellPython

Categories

Site Reliability
PowerPlan, Inc.

About PowerPlan, Inc.

201-500 employees
Contact me