
Reliability Engineer
The Hartford1 month ago
Chicago, IL, USA +3 moreMid Level
Base Salary
$91k - $137k/yr
Responsibilities
- Instrument applications and technology stacks to generate metrics for availability, performance, quality, currency, and resiliency.
- Develop tooling, alerts, automation, and response mechanisms to prevent, detect, mitigate, and resolve reliability risks.
- Improve delivery flows, CI/CD automation, deployment processes, cost efficiency, and self-healing capabilities while adhering to technology standards.
- Coordinate triage and restoration of high-impact incidents and improve monitoring, intelligent incident routing, and automated service restoration.
- Maintain continuity of Hartford and third-party assets and keep IT application and infrastructure metadata repositories current.
- Research and implement AI-based anomaly detection, AI-powered troubleshooting copilots, LLM-driven operational assistants, and AI/ML runbooks.
Requirements
- Bachelor's or master's degree in Computer Science, Engineering, or a related field.
- At least 3 years of experience in Infrastructure Engineering, Site Reliability Engineering, or DevOps.
- Hands-on experience with Splunk, Dynatrace, and CloudWatch, plus deep knowledge of Terraform and CloudFormation.
- Experience optimizing CI/CD pipelines, automating deployments, and applying DevSecOps practices.
- Expertise with AWS and Kubernetes-based microservices environments.
- Strong proficiency in Python and Java for infrastructure automation and tooling development.
- Experience with AI/ML frameworks for observability, predictive failure detection, and AI-driven troubleshooting is desirable.
- Experience with Oracle and SQL Server relational databases; open-source database knowledge is beneficial.
- Experience working within Agile frameworks and strong analytical, problem-solving, and interpersonal skills.
- Candidates must be authorized to work in the US without company sponsorship.
Benefits
- Hybrid schedule requiring work in an office three days per week, Tuesday through Thursday, in Columbus, Ohio; Chicago, Illinois; Hartford, Connecticut; or Charlotte, North Carolina.
- The compensation package may include short-term or annual bonuses, long-term incentives, and on-the-spot recognition.
- The company will not support the STEM OPT I-983 Training Plan endorsement for this position.
Tech Stack
Categories
DevOpsSite Reliability