
Agentic AI Evaluation Engineer
Ernst and Young18 days ago
Kolkāta, IndiaSenior
Responsibilities
- Define and operationalize evaluation strategies for GenAI, RAG, agentic, summarization, extraction, drafting, and multi-step workflow use cases.
- Translate business use cases into evaluation plans covering scope, datasets, metrics, red-team scenarios, thresholds, and reporting requirements.
- Create reusable evaluation templates, test-case libraries, scoring rubrics, dataset requirements, and reporting formats.
- Build Python evaluation harnesses and pipelines for quality, grounding, faithfulness, agent behavior, latency, cost, throughput, stability, and recovery metrics.
- Combine LLM-as-judge and human evaluation using calibrated rubrics, sampling plans, and agreement checks.
- Conduct structured red teaming and adversarial testing for prompt injection, jailbreaks, data leakage, insecure outputs, tool misuse, retrieval poisoning, and denial-of-service patterns.
- Integrate evaluations into development workflows through regression gates, CI checks, and benchmark comparisons across models, prompts, tools, and retrievers.
- Recommend controls including input validation, retrieval filtering, tool sandboxing, least-privilege access, guardrails, refusal logic, output encoding, and monitoring alerts.
- Produce auditable evaluation reports with methodology, datasets, metrics, thresholds, results, risk assessments, and recommended control actions.
- Present residual risk, limitations, and go/no-go recommendations to technical and non-technical stakeholders.
Requirements
- Bachelor’s or master’s degree in data science, statistics, engineering, operational research, or a related field with a strong focus on modern data architectures, processes, and environments.
- 4–7+ years of relevant experience in ML/AI/GenAI/agentic engineering, NLP/LLMs, evaluation engineering, applied research, security testing, red teaming, or evaluation-harness development.
- Strong hands-on Python experience for data processing, metric computation, orchestration, and reporting pipelines.
- Understanding of GenAI architectures including RAG, embeddings, vector search, prompt orchestration, tool calling, multi-agent systems, memory, and routing.
- Experience designing evaluation metrics and methods, including rubrics, automated scoring, sampling strategies, and regression testing.
- Knowledge of LLM risks and mitigations such as data leakage, hallucinations, prompt injection, unsafe content, and bias.
- Understanding of OWASP Top 10 for LLM Applications and adversarial testing methods including jailbreaks, injection, tool misuse, and retrieval poisoning.
- Familiarity with secure-by-design LLM practices such as least privilege, safe tool invocation, output validation or encoding, and monitoring.
- Experience with RAGAS, DeepEval, LangSmith, Phoenix/Arize, or custom evaluation harnesses.
- Experience with experimentation, A/B testing, baseline comparisons, statistical rigor, structured logging, observability, and tracing.
- Basic DevOps experience with Git, CI/CD, Docker, and reproducible environments.
- Strong written, oral, presentation, facilitation, project management, prioritization, organization, and stakeholder-influence skills.
- Preferred experience in assurance, finance, regulatory, model validation, risk acceptance, audit evidence, responsible AI, multilingual evaluation, or enterprise assistants.
- Preferred familiarity with NIST AI RMF, ISO/IEC 42001, EU AI Act concepts, and Azure OpenAI, Azure AI Search, Function Apps, App Insights, and Key Vault.
Benefits
- Support, coaching, and feedback from experienced colleagues.
- Opportunities to develop new skills and progress your career.
- Individual progression planning, education, coaching, and practical experience.
- Interdisciplinary work with knowledge exchange and challenging client assignments.
- Freedom and flexibility to handle the role in a way that is right for you.
- Global collaboration with EY GDS Assurance practices and teams across more than 150 countries.
About Ernst and Young
Ernst & Young (EY) provides audit/assurance, tax, consulting, strategy and transactions services to enterprises, financial institutions, and public‑sector clients. Structured as a global network of partner‑owned member firms, it sells professional services on a fee basis, including a dedicated Financial Services Organization for banking, insurance, and capital markets. Headquartered in London, EY was formed in 1989 from the merger of Ernst & Whinney and Arthur Young, and operates in 150+ countries.