1 day ago
Remote, United StatesSenior
Base Salary
$106k - $149k/yr
Responsibilities
- Own and improve the reliability, availability, performance, and operational health of production data platforms and pipelines.
- Monitor pipeline execution, dependencies, failures, delays, recovery, and downstream impact.
- Troubleshoot production issues across AWS, data platforms, pipelines, and supporting services.
- Design and improve dashboards, alerting, logging, and service-health indicators using Datadog and Splunk.
- Lead or participate in incident response, service restoration, root-cause analysis, and corrective actions.
- Use Python to automate operational tasks, monitoring, health checks, remediation, and repetitive support activities.
- Apply SRE practices including service-level objectives, incident management, capacity planning, resilience, disaster recovery, and operational toil reduction.
- Use SQL for production troubleshooting, validation, and investigation.
- Read, review, and troubleshoot Terraform-managed infrastructure and partner with Cloud or Platform Engineering on deeper infrastructure changes.
- Drive sustainable improvements to recurring operational issues, performance, capacity, resilience, and cloud-resource cost efficiency.
- Develop operational runbooks and recovery procedures.
- Participate in the PagerDuty on-call rotation, including occasional weekend coverage.
Requirements
- At least 5 years of hands-on AWS experience supporting production environments.
- Experience in Site Reliability Engineering, Production Engineering, Platform Engineering, or Cloud Reliability Engineering.
- Strong practical understanding and implementation of SRE principles and production operations.
- Experience supporting Amazon Redshift or similar enterprise data platforms in production.
- Experience owning or supporting the reliability and operations of production data pipelines or data-intensive services.
- Hands-on experience with enterprise observability platforms such as Datadog and Splunk.
- Strong Python programming skills for automation and operational tooling.
- Working knowledge of SQL for production investigation and troubleshooting.
- Strong understanding of Terraform and Infrastructure as Code, including the ability to read, review, and troubleshoot existing IaC.
- Experience with incident management, root-cause analysis, and preventive and corrective actions.
- Hands-on production experience with Snowflake is preferred.
- Experience with multiple cloud-based data platforms, large-scale observability, AIOps, intelligent automation, agentic operations, AI-assisted operations, Amazon CloudWatch, pipeline orchestration, failure recovery, cloud-resource utilization, cost visibility, or maturing SRE practices is preferred.
Benefits
- Medical, dental, vision, and life insurance.
- 401(k) retirement plan with company matching contributions up to 6%, potential discretionary contributions, financial advisory services, and a broad investment lineup.
- Tuition reimbursement up to $5,250 per year.
- Generous paid time off upon hire, including a paid time off program, ten paid company holidays, and three floating holidays annually.
- 16 hours of paid volunteer time per calendar year.
- Paid parental leave, paid short- and long-term disability, and Family and Medical Leave programs.
- Business Resource Groups open to all employees.
- Professional office environment with a business-casual dress policy and the option to wear jeans.
- PagerDuty on-call rotation includes occasional weekend coverage.
- Remote and hybrid employees must provide reliable high-speed wired internet and a suitable home workspace; required computer equipment will be provided.
