
Senior Principal SRE Engineering
Eli Lilly and Company1 day ago
Indianapolis, IN, USAStaff+
Base Salary
$129k - $231k/yr
Responsibilities
- Define SLOs and SLIs across applications based on risk and business impact.
- Govern error-budget burn, burn-rate alerting, reliability reporting, and operating cadences with product and application owners.
- Establish observability, instrumentation, infrastructure-as-code, continuous-delivery hardening, and deployment-safety standards.
- Design and validate self-healing runbooks, graceful degradation, circuit breakers, and automated remediation patterns.
- Lead high-severity incidents as incident commander and ensure postmortems produce durable engineering work.
- Partner with operations and agentic automation engineering teams on safe remediations, confidence thresholds, guardrails, and human-in-the-loop boundaries.
- Ensure reliability practices support regulatory requirements, auditability, security, and validated environments.
- Mentor senior reliability engineers and influence engineering, product, and platform leaders through technical standards and expertise.
Requirements
- Bachelor’s degree in Computer Science, Information Technology, or a related technical engineering discipline.
- At least 7 years of progressive engineering experience, including at least 6 years as a Site Reliability Engineer, Production Engineer, or equivalent.
- Experience serving as the senior-most reliability engineer for a multi-application production estate.
- Hands-on experience designing and operating enterprise-scale SRE platforms using observability tools such as Splunk, Datadog, New Relic, or Grafana/Prometheus.
- Experience with Terraform, CI/CD pipeline hardening, Kubernetes-based container platforms, and production workloads on AWS, Azure, or GCP.
- Experience with SLO/SLI frameworks, error-budget policies, burn-rate alerting, incident command, blameless postmortems, and durable remediation work.
- Experience designing self-healing automation and operating chaos engineering or resilience-testing programs using tools such as AWS Fault Injection Service, Gremlin, or LitmusChaos.
- Deep AWS experience with EKS, ECS, Lambda, CloudWatch, X-Ray, Systems Manager, and Route 53, plus familiarity with the AWS Well-Architected Reliability Pillar.
- Experience operating in regulated or audited environments such as GxP, SOX, HIPAA, PCI, pharma, healthcare, or financial services.
- Experience with AIOps or agent-assisted operations, including guardrails, confidence thresholds, and human-in-the-loop boundaries.
- Demonstrated technical leadership through mentoring senior individual contributors, influencing without formal authority, and authoring engineering standards.
- Experience building a reliability practice from a small founding team is preferred.
Benefits
- Full-time employees may be eligible for a company bonus based on company and individual performance.
- Benefits may include a 401(k), pension, vacation, medical, dental, vision, prescription drug, flexible spending, life insurance, leave, well-being, fitness, employee assistance, and employee club benefits.
- Flexible work hours may be required, including some weekends and holidays to support continuous operations.
- The role is a senior individual contributor position and does not manage people.
Tech Stack
Categories
Site Reliability
About Eli Lilly and Company
Eli Lilly and Company researches, develops, manufactures, and markets prescription medicines and biologics for patients, sold through healthcare providers, pharmacies, and payers. Core therapeutic areas include diabetes, oncology, immunology, neuroscience, and obesity. Founded in 1876 and headquartered in Indianapolis, it is a public company (NYSE: LLY) with global operations and is building a new advanced manufacturing site for gene therapy and other modalities in Lebanon, Indiana.