about 1 month ago
Responsibilities
- Develop a deep understanding of Tekion’s AI agents, ML models, and domain-specific quality dimensions.
- Design, enhance, and own shared AI evaluation infrastructure used across ML teams.
- Create, curate, and maintain evaluation datasets and golden or ground-truth sets.
- Define metrics for accuracy, relevance, faithfulness or groundedness, consistency, safety, and task success.
- Build automated scoring pipelines using LLM-as-judge, rubric-based, and reference-based evaluation methods.
- Validate user intents and response accuracy for AI capabilities such as the Analytics Agent.
- Identify hallucinations, unsafe or biased outputs, and edge cases through targeted evaluation suites.
- Build offline pre-release benchmarks and online production monitoring, including A/B testing and drift detection.
- Establish evaluation gates in CI/CD for model, prompt, and data changes.
- Develop dashboards and reporting that make AI quality actionable for ML and product teams.
- Use AI and LLMs for automated judging, synthetic data generation, and evaluation tooling.
- Champion evaluation and responsible-AI quality practices across the organization.
Requirements
- 5–8 years of experience in SDET, quality engineering, ML engineering, or data science, with hands-on experience building evaluation or measurement systems, or a strong SDET background with deep LLM/ML fluency.
- Strong Python programming skills for building robust, reusable evaluation pipelines and tooling.
- Deep understanding of ML/LLM evaluation, benchmark design, and evaluating non-deterministic systems.
- Hands-on experience with LLM-as-judge, rubric-based scoring, or human-in-the-loop evaluation.
- Understanding of prompting, RAG, embeddings, tool use, and generative failure modes such as hallucination, drift, prompt sensitivity, and bias.
- Experience designing and curating datasets, including labeling or annotation strategy and data quality.
- Strong statistical intuition for interpreting evaluation results and significance.
- Excellent communication skills for translating quality signals into decisions for ML and product teams.
- Preferred experience with evaluation tools such as Ragas, DeepEval, LangSmith, TruLens, Promptfoo, HELM, or provider evaluation suites.
- Preferred experience building online evaluation, guardrails, production model monitoring, responsible AI or safety evaluation, red-teaming, experiment tracking with MLflow or Weights & Biases, A/B testing, or evaluation platforms for multiple teams.