23 days ago
Responsibilities
- Build automated suites to detect hallucinations, bias, toxicity, prompt injection, and other risks in LLM-powered products.
- Evaluate RAG systems for context relevance, groundedness, and answer faithfulness and test multi-agent tool calling, reasoning, memory, and autonomous loops.
- Create automated conversation simulations and prompt-regression frameworks across intents, edge cases, multi-turn dialogs, and parameter changes.
- Statistically validate AI outputs and audit data ingestion, transformation, feature-store pipelines, synthetic datasets, vector indexing, embeddings, and retrieval latency.
- Maintain automated tracking for precision, recall, F1, ROC-AUC, deep-learning loss curves, data drift, and concept drift.
- Build scalable automation for APIs, backend services, and model endpoints.
- Integrate AI evaluation and data-quality suites into MLOps and CI/CD pipelines so quality failures block releases.
- Define and track AI quality KPIs and communicate release readiness to engineering and product teams.
Requirements
- 5+ years of experience in an SDET role.
- Expert-level Python experience for test automation, evaluation pipelines, and data analysis.
- SQL experience for data-output validation, ground-truth queries, and pipeline quality checks.
- Experience with LLM evaluation frameworks such as RAGAS, TruLens, DeepEval, or Promptflow.
- Experience with LangChain, LangSmith, or LlamaIndex for agent workflow testing, prompt tracing, and response debugging.
- Experience testing OpenAI, Anthropic, or Hugging Face APIs and vector databases.
- Experience with Pandas, NumPy, and Pytest for statistical analysis, data profiling, schema validation, and reusable automation suites.
- Experience with Postman, REST Assured, or Requests for API contract and integration testing.
- Experience with MLflow and Docker, GitHub Actions, or Jenkins for model tracking, test environments, and pipeline automation.
- Experience with Grafana, Kibana, or OpenTelemetry for monitoring, drift detection, and distributed tracing.
- Preferred experience includes AWS Bedrock, Azure OpenAI, or GCP Vertex AI.
- Preferred experience includes Kubeflow, Weights & Biases, or Feast feature stores.
- Preferred familiarity with Scikit-learn, TensorFlow, or PyTorch.
- Preferred experience with Kubernetes, Terraform, Playwright, Cypress, Locust, or JMeter.
- Preferred knowledge of statistical hypothesis testing, synthetic data generation, and evaluation dataset creation.
Benefits
- Opportunity to work on a global, cloud-native, AI-driven automotive platform.
- Innovative, collaborative, and fast-paced culture.
- Current Tekion employees should apply via the internal job board in Ashby effective 4 August 2026.