about 1 month ago
Responsibilities
- Develop expertise in Tekion’s AI agents, ML models, business domains, and supporting data.
- Own end-to-end quality for ML models, AI agents, and AI-powered features.
- Design and implement automated testing for model inference, prompt pipelines, agentic workflows, RAG/retrieval systems, and serving APIs.
- Create evaluation datasets and ground-truth sets, measure response quality, identify hallucinations and unsafe outputs, and build automated scoring pipelines.
- Use AI and LLMs to generate test cases, synthetic data, and automation scaffolding.
- Build regression suites to detect quality drift across model, prompt, embedding, and training-data changes.
- Validate ML and agent integration into ARC products, including end-to-end flows and fallback behavior.
- Perform root-cause analysis for production model and behavioral issues and partner with ML Engineers on preventive improvements.
- Define quality metrics, guardrails, and automated monitoring for model degradation.
- Champion quality and evaluation engineering practices across the ML Engineering organization.
Requirements
- 5–8 years of experience in SDET, quality engineering, or software engineering with a strong track record building test automation frameworks.
- Strong programming skills in Python and/or Java with the ability to write production-quality automation code.
- Hands-on experience testing data-intensive or ML/AI systems, or a strong software-testing background with demonstrated ML/LLM fluency.
- Understanding of model training and inference, evaluation metrics such as precision, recall, and F1, and the non-deterministic nature of model outputs.
- Experience designing evaluation frameworks or working with evaluation datasets, benchmarks, or LLM-as-judge approaches.
- Familiarity with prompting, RAG, embeddings, vector search, hallucinations, drift, and prompt sensitivity.
- Experience with CI/CD, test orchestration, and API/integration testing.
- Strong analytical, debugging, root-cause-analysis, collaboration, and communication skills.
- Nice-to-have experience with ML/evaluation tools including MLflow, Weights & Biases, LangSmith, Ragas, DeepEval, or provider evaluation suites.
- Nice-to-have experience with AWS, GCP, or Azure; Docker and Kubernetes; model observability, drift detection, production ML monitoring, and agentic or multi-step LLM workflows.