Eli Lilly and Company

Principal SRE Engineer

Eli Lilly and Company
Apply
1 day ago
Indianapolis, IN, USAStaff+

Base Salary

$66k - $158k/yr

Responsibilities

  • Design and build self-healing automation including circuit breakers, graceful degradation, and automated remediation.
  • Run resilience and chaos testing to validate reliability patterns before production use.
  • Author, validate, maintain, and retire remediation runbooks covering execution order, rollback steps, and exception handling.
  • Define criteria for moving remediation from human-executed to agent-assisted to autonomous operation.
  • Lead or contribute to root-cause analyses, blameless postmortems, durable fixes, and high-severity incident response, including incident command.
  • Partner with agentic automation engineering and Operations teams on safe remediation, confidence thresholds, human-in-the-loop boundaries, and outcome validation.
  • Ensure automation and runbooks meet change-control, audit, and validated-environment standards.
  • Mentor reliability and automation engineers and contribute proven patterns to the broader reliability practice.

Requirements

  • Bachelor’s degree in Computer Science, Information Technology, or a related technical engineering discipline.
  • At least 5 years of progressive engineering experience, including meaningful Site Reliability Engineer, Production Engineer, or equivalent experience.
  • Hands-on ownership of self-healing automation or runbook-driven remediation for a multi-application production estate.
  • Production reliability experience in a regulated or audited environment such as GxP, SOX, HIPAA, PCI, or equivalent.
  • Experience authoring and validating runbooks with safe execution order, rollback steps, and exception handling.
  • Hands-on experience with enterprise-scale SRE platforms, observability solutions, Terraform, CI/CD pipeline hardening, Kubernetes-based container platforms, and AWS, Azure, or GCP workloads.
  • Experience designing self-healing patterns and validating them before production use.
  • Preferred experience with chaos engineering or resilience testing using AWS Fault Injection Service, Gremlin, LitmusChaos, or equivalent.
  • Preferred deep AWS experience with EKS, ECS, Lambda, CloudWatch, X-Ray, Systems Manager, Route 53, and the AWS Well-Architected Reliability Pillar.
  • Preferred experience with AIOps or agent-assisted operations, autonomy criteria, runbook libraries, regulated industries, root-cause analysis, and blameless postmortems.
  • Must be authorized to work full-time in the United States; Lilly will not sponsor work authorization or visas for this role.

Benefits

  • Eligible full-time employees may participate in a company bonus program.
  • Benefits may include a 401(k), pension, vacation, medical, dental, vision, prescription drug, flexible spending, life insurance, leave, well-being, fitness, employee assistance, and employee club benefits.
  • Flexible work hours may be required, including non-standard hours, weekends, and holidays, to support continuous operations.

Tech Stack

Categories

Site Reliability
Eli Lilly and Company

About Eli Lilly and Company

10,000+ employees

Eli Lilly and Company researches, develops, manufactures, and markets prescription medicines and biologics for patients, sold through healthcare providers, pharmacies, and payers. Core therapeutic areas include diabetes, oncology, immunology, neuroscience, and obesity. Founded in 1876 and headquartered in Indianapolis, it is a public company (NYSE: LLY) with global operations and is building a new advanced manufacturing site for gene therapy and other modalities in Lebanon, Indiana.

Contact me