
Generative AI Cloud Operations Engineer - Evinova
AstraZeneca2 hours ago
Gaithersburg, MD, USAMid Level
H1B sponsor
Base Salary
$145k - $185k/yr
Responsibilities
- Develop and manage Generative AI operations systems for clinical trial design, planning, and operational optimization.
- Design, implement, deploy, monitor, and operate resilient cloud-native LLM agent and multi-agent capabilities in production.
- Integrate LLM proxies and routers such as LiteLLM Proxy/Router and optimize and scale RAG pipelines.
- Implement platform-level monitoring for token usage, latency, response quality, hallucination detection, reliability, performance, and cost.
- Partner with AI engineers and data scientists to transition research projects into production-grade agentic Generative AI capabilities.
- Use and promote frameworks and tools including LangChain, LangGraph, Google ADK, Langfuse, DSPy, Arize Phoenix, Pinecone, Weaviate, Splunk, Grafana, Prometheus, and Xray.
- Manage AWS cloud resources, infrastructure-as-code, CI/CD pipelines, and cloud-native AI/ML workloads.
- Improve system scalability, reliability, performance, governance, operational processes, and risk mitigation.
- Provide reliable access to large language models through Vertex AI, Azure Foundry, OpenAI, Anthropic, and other foundation model platforms.
- Document technical solutions, operational processes, and system performance while collaborating with cross-functional stakeholders.
Requirements
- High school diploma or GED required.
- At least 2 years of hands-on experience deploying, operating, and maintaining Generative AI agents, workflows, or applications in production.
- Strong understanding of production Generative AI concerns including reliability, scalability, latency, cost optimization, observability, evaluation, and model performance.
- Hands-on experience deploying agentic AI solutions with LangChain, LangGraph, LlamaIndex, Google ADK, Strands Agents, or similar frameworks.
- Strong experience with LLM evaluation and observability platforms such as Arize Phoenix, Langfuse, Braintrust, Freeplay, or comparable tools.
- Strong software engineering skills in Python and/or TypeScript and experience building production-quality systems.
- Deep expertise with AWS cloud services and cloud-native AI/ML workload operations.
- Strong infrastructure-as-code experience, including AWS CDK with Python and/or TypeScript.
- Experience with Docker and Kubernetes.
- Understanding of the data science and machine learning lifecycle, including moving models and AI capabilities from experimentation through production and ongoing operations.
- Experience operationalizing RAG pipelines, LLM applications, or multi-agent systems in production is strongly preferred.
- Ability to stay current with evolving Generative AI models, frameworks, tooling, evaluation techniques, and engineering practices.
- Ability to collaborate with AI/ML engineers, data scientists, software engineers, product teams, and other stakeholders.
- Strong written and verbal communication skills for documenting technical and operational information.
Benefits
- Annual base pay ranges from $145,000 to $185,000 USD, with an additional short-term incentive bonus opportunity and potential equity-based long-term incentive participation.
- Benefits include a qualified retirement program with a 401(k) plan, paid vacation and holidays, paid leaves, and medical, prescription drug, dental, and vision coverage.
- Hybrid work arrangement requiring employees to work from the office three days per week.
- Equal opportunity employer with disability accommodations available throughout the recruitment, assessment, and selection process.
- Position is an at-will employment role.
- Applications are accepted through September 10, 2026, with the posting dated August 17, 2026.
Tech Stack
Categories
About AstraZeneca
We're transforming the future of healthcare by unlocking the power of what science can do for people, society and the planet. For more information, visit www.astrazeneca.com. Community Guidelines: bit.ly/2MgAcio