Empower

Senior Engineer Site Reliability - Data Operations

Empower
Apply
1 day ago
Remote, United StatesSenior

Base Salary

$106k - $149k/yr

Responsibilities

  • Own and improve the reliability, availability, performance, and operational health of production data platforms and pipelines.
  • Monitor pipeline execution, dependencies, failures, delays, recovery, and downstream impact.
  • Troubleshoot production issues across AWS, data platforms, pipelines, and supporting services.
  • Design and improve dashboards, alerting, logging, and service-health indicators using Datadog and Splunk.
  • Lead or participate in incident response, service restoration, root-cause analysis, and corrective actions.
  • Use Python to automate operational tasks, monitoring, health checks, remediation, and repetitive support activities.
  • Apply SRE practices including service-level objectives, incident management, capacity planning, resilience, disaster recovery, and operational toil reduction.
  • Use SQL for production troubleshooting, validation, and investigation.
  • Read, review, and troubleshoot Terraform-managed infrastructure and partner with Cloud or Platform Engineering on deeper infrastructure changes.
  • Drive sustainable improvements to recurring operational issues, performance, capacity, resilience, and cloud-resource cost efficiency.
  • Develop operational runbooks and recovery procedures.
  • Participate in the PagerDuty on-call rotation, including occasional weekend coverage.

Requirements

  • At least 5 years of hands-on AWS experience supporting production environments.
  • Experience in Site Reliability Engineering, Production Engineering, Platform Engineering, or Cloud Reliability Engineering.
  • Strong practical understanding and implementation of SRE principles and production operations.
  • Experience supporting Amazon Redshift or similar enterprise data platforms in production.
  • Experience owning or supporting the reliability and operations of production data pipelines or data-intensive services.
  • Hands-on experience with enterprise observability platforms such as Datadog and Splunk.
  • Strong Python programming skills for automation and operational tooling.
  • Working knowledge of SQL for production investigation and troubleshooting.
  • Strong understanding of Terraform and Infrastructure as Code, including the ability to read, review, and troubleshoot existing IaC.
  • Experience with incident management, root-cause analysis, and preventive and corrective actions.
  • Hands-on production experience with Snowflake is preferred.
  • Experience with multiple cloud-based data platforms, large-scale observability, AIOps, intelligent automation, agentic operations, AI-assisted operations, Amazon CloudWatch, pipeline orchestration, failure recovery, cloud-resource utilization, cost visibility, or maturing SRE practices is preferred.

Benefits

  • Medical, dental, vision, and life insurance.
  • 401(k) retirement plan with company matching contributions up to 6%, potential discretionary contributions, financial advisory services, and a broad investment lineup.
  • Tuition reimbursement up to $5,250 per year.
  • Generous paid time off upon hire, including a paid time off program, ten paid company holidays, and three floating holidays annually.
  • 16 hours of paid volunteer time per calendar year.
  • Paid parental leave, paid short- and long-term disability, and Family and Medical Leave programs.
  • Business Resource Groups open to all employees.
  • Professional office environment with a business-casual dress policy and the option to wear jeans.
  • PagerDuty on-call rotation includes occasional weekend coverage.
  • Remote and hybrid employees must provide reliable high-speed wired internet and a suitable home workspace; required computer equipment will be provided.

Tech Stack

Amazon RedshiftAWSDatadogPythonSnowflakeSplunkSQLTerraform

Categories

Site Reliability
Empower

About Empower

10,000+ employees
Contact me