GrepJob
Tekion

Staff Software Development Test Engineer - AI Evaluation

Tekion
Apply
about 1 month ago
Bengaluru, IndiaStaff+
H1B Sponsor

Responsibilities

  • Develop a deep understanding of Tekion’s AI agents, ML models, and domain-specific quality dimensions.
  • Design, enhance, and own shared AI evaluation infrastructure used across ML teams.
  • Create, curate, and maintain evaluation datasets and golden or ground-truth sets.
  • Define metrics for accuracy, relevance, faithfulness or groundedness, consistency, safety, and task success.
  • Build automated scoring pipelines using LLM-as-judge, rubric-based, and reference-based evaluation methods.
  • Validate user intents and response accuracy for AI capabilities such as the Analytics Agent.
  • Identify hallucinations, unsafe or biased outputs, and edge cases through targeted evaluation suites.
  • Build offline pre-release benchmarks and online production monitoring, including A/B testing and drift detection.
  • Establish evaluation gates in CI/CD for model, prompt, and data changes.
  • Develop dashboards and reporting that make AI quality actionable for ML and product teams.
  • Use AI and LLMs for automated judging, synthetic data generation, and evaluation tooling.
  • Champion evaluation and responsible-AI quality practices across the organization.

Requirements

  • 5–8 years of experience in SDET, quality engineering, ML engineering, or data science, with hands-on experience building evaluation or measurement systems, or a strong SDET background with deep LLM/ML fluency.
  • Strong Python programming skills for building robust, reusable evaluation pipelines and tooling.
  • Deep understanding of ML/LLM evaluation, benchmark design, and evaluating non-deterministic systems.
  • Hands-on experience with LLM-as-judge, rubric-based scoring, or human-in-the-loop evaluation.
  • Understanding of prompting, RAG, embeddings, tool use, and generative failure modes such as hallucination, drift, prompt sensitivity, and bias.
  • Experience designing and curating datasets, including labeling or annotation strategy and data quality.
  • Strong statistical intuition for interpreting evaluation results and significance.
  • Excellent communication skills for translating quality signals into decisions for ML and product teams.
  • Preferred experience with evaluation tools such as Ragas, DeepEval, LangSmith, TruLens, Promptfoo, HELM, or provider evaluation suites.
  • Preferred experience building online evaluation, guardrails, production model monitoring, responsible AI or safety evaluation, red-teaming, experiment tracking with MLflow or Weights & Biases, A/B testing, or evaluation platforms for multiple teams.

Tech Stack

HelmMLflowPython

Categories