
Senior Principal Agentic Engineer
Eli Lilly and Company3 days ago
Hyderābād, IndiaStaff+
Responsibilities
- Architect and build the agentic resolution platform from intake and classification through action execution, verification, and human-in-the-loop fallback.
- Develop agent runtimes and orchestration capabilities involving state, memory, tool integration, multi-agent coordination, and confidence-based handoffs.
- Design production LLM patterns including prompt engineering, retrieval-augmented generation, structured outputs, multi-model routing, hybrid retrieval, and evaluation frameworks.
- Architect and operate Kubernetes-based cloud-native infrastructure, write production platform services and tooling, and maintain Terraform-based infrastructure as code.
- Build CI/CD, observability, reliability, SLO, SLI, security, and developer-experience capabilities for platform and agentic systems.
- Set engineering standards, conduct architectural reviews, mentor senior engineers, and partner with reliability teams and senior architects.
Requirements
- 12+ years of progressive technology experience with hands-on architecture and delivery of automation, AIOps, agentic systems, or cloud-native platforms at enterprise scale.
- Demonstrated ownership of measurable production deflection or operational-toil-reduction outcomes.
- Deep experience with agent runtimes, multi-agent orchestration, tool integration, memory, state management, and at least one of LangGraph, LangChain, LlamaIndex, or MCP.
- Production experience with prompt engineering, retrieval-augmented generation, structured outputs, and multi-model routing.
- Deep cloud-native experience with Kubernetes at scale, container workloads, Terraform, and at least one major cloud AI stack.
- Expert-level Python experience for production services; Go experience is welcome.
- Strong CI/CD, model and agent deployment, versioning, canary or blue-green rollout, evaluation gate, and rollback experience.
- Experience with AI observability using metrics, logs, traces, and agent-decision telemetry, including Prometheus, Grafana, OpenTelemetry, and an enterprise observability platform.
- Experience with evaluation, guardrails, drift detection, human-in-the-loop verification, feedback loops, security controls, and regulated environments.
- Bachelor’s degree or higher in Computer Science, Information Technology, or a closely related field.
- Preferred qualifications include operational ML, AI observability tooling, ServiceNow and Microsoft Teams integrations, vector databases, internal developer platforms, REST/gRPC APIs, databases, message queues, distributed systems, policy-as-code, FinOps, regulated industries, capability building, and React or TypeScript.
Benefits
- Onsite work in Hyderabad with flexible hours and different shifts to support global delivery partners.
- Some weekend and holiday work may be required for continuous operations, with applicable benefit adjustments for non-standard hours.
- Disability accommodations and equal-opportunity employment support are available.