3 months ago
Bengaluru, IndiaStaff+
Responsibilities
- Develop understanding of Tekion’s AI agents, ML models, and domain-specific quality dimensions.
- Design, enhance, and own shared AI evaluation infrastructure used across ML teams.
- Create, curate, and maintain evaluation datasets and golden or ground-truth sets.
- Define metrics for accuracy, relevance, faithfulness or groundedness, consistency, safety, and task success.
- Build automated scoring pipelines using LLM-as-judge, rubric-based, and reference-based evaluation methods.
- Validate user intents and measure response accuracy and consistency for AI-powered capabilities such as the Analytics Agent.
- Identify hallucinations, unsafe or biased outputs, and edge cases, and create targeted evaluation suites.
- Build offline pre-release benchmarks and online production evaluation, monitoring, A/B testing, and drift detection.
- Establish evaluation gates in CI/CD for model, prompt, and data changes.
- Develop dashboards and reporting that make AI quality visible and actionable.
- Use AI and LLMs to build automated judges, synthetic datasets, and evaluation tooling.
- Promote evaluation and responsible-AI quality practices across the organization.
Requirements
- 5–8 years of experience in SDET, quality engineering, ML engineering, or data science, with hands-on experience building evaluation or measurement systems, or a strong SDET background with deep LLM/ML fluency.
- Strong Python programming skills for building robust, reusable evaluation pipelines and tooling.
- Deep understanding of ML/LLM evaluation, benchmark design, and evaluating non-deterministic systems.
- Hands-on experience with LLM-as-judge, rubric-based scoring, or human-in-the-loop evaluation.
- Understanding of prompting, RAG, embeddings, tool use, hallucinations, drift, prompt sensitivity, and bias.
- Experience designing and curating datasets, including labeling or annotation strategy and data quality.
- Strong statistical intuition for interpreting evaluation results and significance.
- Excellent communication skills for translating quality signals into decisions for ML and product teams.
- Preferred experience with Ragas, DeepEval, LangSmith, TruLens, Promptfoo, HELM, or provider evaluation suites.
- Preferred experience building online evaluation, guardrails, production model monitoring, responsible AI or safety evaluation, red-teaming, experiment tracking, and A/B testing.
- Prior experience establishing evaluation as a platform capability for multiple teams is preferred.
Tech Stack
Categories
About Tekion
Tekion builds an AI-native, cloud platform for automotive retail that unifies dealers, OEMs, and partners. Its products include Automotive Retail Cloud (a dealership management system for retailers), Automotive Enterprise Cloud for manufacturers, and Automotive Partner Cloud for integrations, delivered as subscription software. Privately held and headquartered in Pleasanton, California, Tekion raised private equity funding in 2024.
