23 days ago
Responsibilities
- Build automated suites to detect hallucinations, bias, toxicity, prompt injection, and other risks in LLM-powered products.
- Implement evaluations for RAG systems, multi-agent workflows, prompt regressions, conversation simulations, and model-output consistency.
- Validate AI data distributions, precision and recall, error patterns, ingestion and transformation pipelines, feature stores, vector indexing, embeddings, retrieval latency, and synthetic datasets.
- Maintain automated tracking for ML metrics and deep-learning loss curves and monitor live endpoints for data and concept drift.
- Build scalable test automation for APIs, backend services, and model endpoints.
- Integrate AI evaluation and data-quality suites into MLOps and CI/CD pipelines so quality failures can block releases.
- Define AI-quality KPIs and communicate release readiness to engineering and product teams.
Requirements
- 8+ years of experience in an SDET role.
- Expert-level Python experience with test automation, evaluation pipelines, and data analysis using Pandas, NumPy, and Pytest.
- SQL experience for data-output validation, ground-truth querying, and pipeline data-quality checks.
- Experience with LLM evaluation frameworks such as RAGAS, TruLens, DeepEval, and Promptflow.
- Experience with LangChain, LangSmith, or LlamaIndex for agent-workflow testing, prompt tracing, and LLM response debugging.
- Experience testing OpenAI, Anthropic, or Hugging Face APIs and vector databases.
- Experience with API and automation tools including Pytest, Postman, REST Assured, and Requests.
- Experience with MLflow, Docker, GitHub Actions, and Jenkins for MLOps and CI/CD automation.
- Experience with Grafana, Kibana, or OpenTelemetry for observability and distributed tracing.
- Preferred familiarity with AWS Bedrock, Azure OpenAI, GCP Vertex AI, Kubeflow, Weights & Biases, Feast, Scikit-learn, TensorFlow, PyTorch, Kubernetes, Terraform, Playwright, Cypress, Locust, and JMeter.
- Preferred knowledge of statistical hypothesis testing and synthetic data generation for AI evaluation.