
Principal Reliability Engineer - EDS
The Hartford2 hours ago
Remote, United States or Tokyo, JapanStaff+
Base Salary
$153k - $229k/yr
Responsibilities
- Define the Reliability Engineering strategy, roadmaps, operating models, and architectural patterns for enterprise data platforms and products.
- Serve as the highest-level technical escalation point for systemic reliability issues and influence executive stakeholders and engineering leaders.
- Architect reliable, performant, and cost-efficient platforms across AWS and GCP.
- Oversee reliability controls and fail-safe patterns for Snowflake, EMR, Hadoop/Spark clusters, Kubernetes-based container platforms, and mission-critical data systems.
- Establish SLI/SLO frameworks, observability standards, incident response patterns, post-incident reviews, and continuous improvement processes.
- Develop AI-driven automation for anomaly detection, alert correlation, autonomous remediation, predictive capacity management, and intelligent operational tooling.
- Define reliability practices for data products, governed pipelines, real-time systems, and operational analytics platforms.
- Embed resilience patterns such as idempotency, checkpointing, replayability, and disaster recovery into data pipeline architectures.
- Set and enforce standards for infrastructure as code, CI/CD, platform automation, reliability frameworks, operational readiness, and runbook quality.
- Provide technical leadership and mentorship to Staff and Senior Engineers and represent Reliability Engineering in architecture, governance, and executive forums.
Requirements
- 10+ years of experience in data, cloud, platform engineering, site reliability engineering, or large-scale distributed systems, including leadership or technology-leader experience.
- Proficiency with data and cloud platforms and architectural patterns for resilience, networking, security, and distributed data infrastructure.
- Deep experience with Snowflake, EMR, Hadoop/Spark, data integration, and cloud-native data ecosystems.
- Experience with Python for large-scale automation, platform tooling, and reliability frameworks.
- Experience with Terraform, CloudFormation, and enterprise CI/CD.
- Experience designing or operating observability stacks such as Prometheus, Grafana, Datadog, Splunk, Dynatrace, or OpenTelemetry.
- Experience with SLI/SLO frameworks, incident response, data quality, lineage, metadata, governance, and pipeline reliability practices.
- Experience applying machine learning, LLMs, prompt engineering, or cloud AI services to operations and reliability improvements.
- Preferred experience in regulated or complex enterprise environments such as financial services, insurance, or healthcare.
- Prior experience as a Senior Staff Engineer, engineering or architecture leader, or similar senior technical role with hands-on experience.
- Strong ability to lead technical strategy, influence senior leaders, mentor engineers, provide architectural guidance, and communicate with executives and cross-functional teams.
- AWS, GCP, Kubernetes, or SRE/DevOps certifications are preferred.
Benefits
- Hybrid schedule requiring work in an office three days per week, Tuesday through Thursday, in Columbus, OH; Chicago, IL; Hartford, CT; or Charlotte, NC.
- The role offers an annualized base pay range of $152,800 to $229,200, plus potential bonuses, long-term incentives, and recognition rewards.
- The company does not support the STEM OPT I-983 Training Plan endorsement for this position.
Tech Stack
Apache HadoopApache SparkAWSGoogle Cloud PlatformGrafanaKubernetesPrometheusPythonSnowflakeSplunkTerraform
Categories
Site Reliability