
Senior Principal Application Support Engineering Lead – AI Products & Agents
Eli Lilly and Company21 days ago
Hyderābād, IndiaStaff+
Responsibilities
- Lead end-to-end application support operations across shifts and time zones, including handoffs, escalation paths, and operational decision frameworks.
- Serve as incident commander for complex, high-impact production incidents and coordinate recovery with engineering, product, platform, security, and vendor teams.
- Own enterprise problem management, root-cause analysis, defect elimination, and measurable corrective actions.
- Define reliability strategy using SRE and production engineering principles, including SLIs, SLOs, error budgets, readiness standards, runbooks, monitoring, and rollback plans.
- Set observability direction across logs, metrics, and traces and drive automation, reusable runbooks, and remediation patterns to reduce toil and MTTR.
- Provide operational oversight for deployments, releases, risk assessment, go/no-go decisions, and post-release validation.
- Ensure compliant, secure, auditable support operations in regulated and validated environments.
- Mentor R2/R3 engineers and influence managers, architects, and engineering leaders through technical expertise and operational outcomes.
Requirements
- 10–14+ years of experience in application support, production engineering, SRE, or software engineering with deep operational ownership.
- Extensive experience leading high-severity incident response and operational execution.
- Deep hands-on troubleshooting experience across distributed applications, integrations, databases, and cloud platforms.
- Strong experience with monitoring, logging, and alerting platforms such as Datadog, Splunk, ELK, AppDynamics, or CloudWatch.
- Advanced scripting and automation skills using tools such as Python and Bash.
- Proven ability to work in regulated enterprise environments and communicate and lead effectively under pressure.
- Experience with Power Automate, LLM models, generative AI, and agentic AI.
- Preferred experience with SLIs, SLOs, error budgets, containers, Kubernetes, CI/CD pipelines, Infrastructure as Code, globally distributed support teams, and enterprise operational standards.
Benefits
- Full-time onsite position in Hyderabad.
- Flexible and non-standard work hours are required for continuous two-shift operations, including possible weekends and holidays.
- Candidates should be open to shifts from 6AM–2PM and 2PM–11PM.
- Appropriate benefits adjustments may be provided for employees working non-standard hours where applicable.
Tech Stack
Categories
Site Reliability