Eli Lilly and Company

Senior Principal SRE Engineering

Eli Lilly and Company
Apply
24 hours ago
Hyderābād, IndiaStaff+

Responsibilities

  • Define SLOs and SLIs across the supported production estate and govern error budgets, burn rates, and reliability reporting.
  • Establish observability, instrumentation, infrastructure-as-code, continuous-delivery, and deployment-safety standards.
  • Convert incidents and root-cause analyses into durable engineering fixes, self-healing runbooks, graceful degradation, and circuit-breaker patterns.
  • Lead high-severity incident response as incident commander and promote blameless postmortems and measurable reduction in recurrence.
  • Partner with agentic automation teams to establish safe remediation guardrails, confidence thresholds, and human-in-the-loop boundaries.
  • Ensure reliability practices meet regulatory, auditability, security, and validated-environment requirements.
  • Mentor senior reliability engineers and influence engineering, product, and platform leaders through technical standards and expertise.

Requirements

  • 14+ years of progressive engineering experience, including at least 6 years as an SRE, Production Engineer, or equivalent, with experience serving as the senior-most reliability engineer for a multi-application production estate.
  • Production reliability experience in a regulated or audited environment such as GxP, SOX, HIPAA, PCI, or equivalent.
  • Hands-on ownership of SLO/SLI frameworks, error-budget policies, burn-rate alerting, and negotiations with product owners.
  • Deep experience with observability, OpenTelemetry, infrastructure-as-code, CI/CD hardening, progressive delivery, Kubernetes, container platforms, and at least one major public cloud.
  • Experience with observability tools such as Splunk, Datadog, New Relic, or Grafana/Prometheus, plus Terraform.
  • Experience leading high-severity incidents as incident commander, running blameless postmortems, and delivering durable engineering improvements.
  • Demonstrated technical leadership through mentoring senior individual contributors, influencing without formal authority, and authoring engineering standards.
  • Bachelor's degree or higher in Computer Science, Information Technology, or a closely related engineering field.
  • Preferred experience includes self-healing automation, chaos engineering, resilience testing, AWS reliability services, AWS Well-Architected Reliability Pillar, AIOps, agent-assisted operations, founding a reliability practice, and regulated-industry experience.

Benefits

  • Onsite role based in Hyderabad.
  • Flexible and non-standard work hours may be required, including shifts, weekends, and holidays, with appropriate benefit adjustments where applicable.
  • Lilly provides workplace accommodation support and promotes an inclusive, non-discriminatory work environment.

Tech Stack

Categories

Site Reliability
Eli Lilly and Company

About Eli Lilly and Company

10,000+ employees

Eli Lilly and Company researches, develops, manufactures, and markets prescription medicines and biologics for patients, sold through healthcare providers, pharmacies, and payers. Core therapeutic areas include diabetes, oncology, immunology, neuroscience, and obesity. Founded in 1876 and headquartered in Indianapolis, it is a public company (NYSE: LLY) with global operations and is building a new advanced manufacturing site for gene therapy and other modalities in Lebanon, Indiana.

Contact me