
Generative AI Cloud Operations Engineer - Evinova
AstraZeneca2 hours ago
Mississauga, CanadaMid Level
Base Salary
$135k - $177k/yr
Responsibilities
- Develop and manage Generative AI operations systems for clinical trial design, planning, and operational optimization.
- Design, implement, deploy, monitor, and operate resilient cloud-based Generative AI agent capabilities.
- Integrate LLM proxies and routers such as LiteLLM Proxy/Router and optimize RAG pipelines for scale.
- Implement platform-level monitoring for token usage, latency, response quality, and hallucination detection.
- Partner with AI engineers and data scientists to move research projects into production-grade agentic AI capabilities.
- Use modern frameworks and tools to design, validate, deploy, and monitor Generative AI agents in production.
- Improve system scalability, reliability, performance, cost efficiency, risk mitigation, and operational processes.
- Provide reliable access to large language models through Vertex AI, Azure Foundry, OpenAI, Anthropic, and other foundation model platforms.
- Ensure predictions and AI capabilities are supported by exploratory data analysis and are interpretable, explainable, safe, and actionable.
- Document technical solutions, operational processes, and system performance while collaborating with cross-functional stakeholders.
Requirements
- High school diploma or GED is required.
- At least 2 years of hands-on experience deploying, operating, and maintaining Generative AI agents, workflows, or applications in production environments.
- Strong understanding of production Generative AI concerns including reliability, scalability, latency, cost optimization, observability, evaluation, and model performance.
- Hands-on experience deploying agentic AI solutions with frameworks such as LangChain, LangGraph, LlamaIndex, Google ADK, or Strands Agents.
- Strong experience with LLM evaluation and observability platforms such as Arize Phoenix, Langfuse, Braintrust, or Freeplay.
- Strong Python and/or TypeScript software engineering skills and experience building production-quality systems.
- Deep expertise with AWS cloud services for cloud-native AI and machine learning workloads.
- Strong infrastructure-as-code experience, including AWS CDK with Python and/or TypeScript.
- Experience with Docker and Kubernetes.
- Understanding of the data science and machine learning lifecycle, including moving models and AI capabilities from experimentation through production and ongoing operations.
- Experience operationalizing RAG pipelines, LLM applications, or multi-agent systems in production is preferred.
- Ability to stay current with Generative AI models, frameworks, tooling, evaluation techniques, and engineering practices.
- Ability to collaborate with AI/ML engineers, data scientists, software engineers, product teams, and other stakeholders.
- Strong written and verbal communication skills.
Benefits
- Hybrid work arrangement with employees expected to work from the office three days per week.
- Permanent positions offer a Flex Benefits and Retirement Savings Program, four weeks of paid vacation, Personal Days, annual variable pay or short-term incentive eligibility, and potential equity-based long-term incentive eligibility.
- Fixed-term contract and temporary positions, excluding student positions, offer a Contract Benefits Program.
- Equal opportunity employer with disability accommodations available upon request.
Tech Stack
Categories
About AstraZeneca
We're transforming the future of healthcare by unlocking the power of what science can do for people, society and the planet. For more information, visit www.astrazeneca.com. Community Guidelines: bit.ly/2MgAcio