
Senior Product Software Engineer - AI Quality Engineering
Wolters Kluwer1 day ago
Base Salary
$83k - $146k/yr
Responsibilities
- Own and implement quality engineering strategy for enterprise generative AI applications, including RAG pipelines, LLM orchestration, agentic workflows, APIs, ingestion pipelines, and AI-driven experiences.
- Architect evaluation harnesses measuring retrieval accuracy, citation correctness, relevance, groundedness, hallucination rate, safety, latency, cost, and end-to-end behavior.
- Design automated regression suites covering prompts, embeddings, chunking strategies, model versions, retrieval configuration, orchestration logic, and application workflows.
- Validate system architecture, service integrations, API behavior, data ingestion, observability, resilience, security controls, and production readiness.
- Lead root-cause analysis for AI quality failures using logs, traces, evaluation outputs, feedback signals, customer scenarios, and acceptance criteria.
- Define and maintain measurable AI release quality gates and advise leaders on risk, readiness, and tradeoffs.
- Develop Python test automation for APIs, services, chat workflows, data pipelines, and AI evaluation workflows within CI/CD pipelines.
- Build observability practices and dashboards tracking model behavior, retrieval performance, operational health, quality trends, and customer-impacting failures.
- Implement responsible AI quality practices, including bias and fairness checks, content safety validation, prompt injection testing, privacy validation, and compliance evidence collection.
- Translate customer workflows and business requirements into testable AI quality criteria with product, UX, subject matter experts, and other stakeholders.
- Mentor engineers and quality team members in AI evaluation, non-deterministic testing, automation, and production-quality engineering practices.
- Stay current with AI quality, LLM evaluation, RAG assessment, and LLMOps practices through research and experimentation.
Requirements
- At least 5 years of experience in quality engineering, SDET, software engineering, test automation, reliability engineering, or equivalent engineering roles.
- At least 2 years of hands-on experience testing, evaluating, or validating AI/ML, LLM, RAG, or other non-deterministic systems in production or production-like environments.
- Strong proficiency with Python test automation frameworks such as Pytest, unittest, or equivalent, including reusable test harnesses and regression suites.
- Experience validating REST APIs, microservices, service integrations, data ingestion pipelines, and end-to-end web or chat workflows.
- Practical understanding of RAG systems, vector search, embeddings, prompt behavior, model evaluation, groundedness, citation quality, and AI failure modes.
- Experience defining quality metrics and success criteria for probabilistic outputs.
- Knowledge of CI/CD pipelines, Git-based workflows, automated regression testing, release gates, and modern engineering practices.
- Experience with observability, logging, tracing, dashboards, and production metrics.
- Ability to decompose complex AI behaviors into testable risks and communicate findings to technical and non-technical stakeholders.
- Strong collaboration skills in Agile environments and ability to influence quality practices across product, design, development, and domain expert teams.
- Commitment to responsible AI practices involving trust, transparency, security, privacy, fairness, and customer impact.
- Preferred experience with LLM observability and evaluation platforms such as LangSmith, Langfuse, TruLens, Arize/Phoenix, or custom tooling.
- Preferred experience with vector databases or search platforms such as Azure AI Search, Pinecone, Weaviate, Elasticsearch, or similar technologies.
- Preferred experience with LangChain, LlamaIndex, Semantic Kernel, Azure OpenAI Service, Azure AI Studio, Azure cloud services, Docker, Kubernetes, Locust, k6, JMeter, Azure Application Insights, Grafana, Prometheus, Datadog, or OpenTelemetry.
- Preferred knowledge of LLMOps, model monitoring, A/B testing, human-in-the-loop evaluation, production feedback loops, and sensitive or regulated data.
- Background in tax, accounting, legal, financial, or other domains where accuracy and explainability are critical.
Benefits
- Medical, dental, and vision plans.
- 401(k), FSA/HSA, commuter benefits, tuition assistance, vacation and sick time, and paid parental leave.
- The role is eligible for a bonus.
- Applicants may be required to appear onsite at a Wolters Kluwer office as part of the recruitment process.
Tech Stack
Categories
About Wolters Kluwer
Wolters Kluwer builds subscription software, data, and workflow solutions for professionals in healthcare, tax and accounting, legal and regulatory, finance and compliance, and ESG. Its portfolio spans SaaS applications and expert content that support compliance, decision-making, and operational processes. Founded in 1836 and headquartered in Alphen aan den Rijn, the Netherlands, it is publicly traded on Euronext Amsterdam (WKL) and reported 2025 annual revenue of €6.1 billion.