2 days ago
Mountain View, CA, USAStaff+
Base Salary
$218k - $271k/yr
Responsibilities
- Own the framework for evaluating, validating, and monitoring AI agents across development, staging, and production.
- Design and maintain evaluation pipelines for LLM outputs, agent behavior, tool use, and multi-turn interactions.
- Instrument agentic systems to detect behavioral drift, regressions, latency issues, hallucinations, tool misuse, and policy violations.
- Lead AI test strategies including red-teaming, golden dataset construction, LLM-as-judge pipelines, and property-based testing.
- Build internal developer tooling, feedback loops, and workflows that accelerate local testing and evaluation runs.
- Establish quality standards, patterns, and education for safely developing and deploying AI-powered features.
- Build and maintain RAG pipelines with strategic data chunking and pre-retrieval data quality gates.
- Partner with Security, Platform, Product, and AI/ML teams to embed quality gates into agent workflows.
- Mentor senior and mid-level engineers on AI evaluation, observability, and testing practices.
Requirements
- Bachelor’s degree in Computer Science, Engineering, or equivalent experience.
- 8+ years building and operating production software systems.
- Demonstrated experience evaluating or testing LLM-powered features or autonomous agents in production.
- Proficiency with AI-assisted development tools such as Claude Code, Cursor, or equivalent.
- Strong backend engineering fundamentals in Python, Java, Go, or equivalent.
- Experience designing test infrastructure, quality gates, or evaluation pipelines at scale.
- Experience improving developer experience through internal tooling, toil reduction, or engineering workflow acceleration.
- Ability to lead cross-team technical initiatives and influence engineering standards.
- Strong written and verbal communication across engineering, product, and leadership.
- Experience building LLM-agent evaluation frameworks, including correctness graders, LLM-as-judge systems, human-in-the-loop evaluations, or benchmark dataset curation.
- Familiarity with agentic frameworks such as Claude API, Anthropic SDK, BrainTrust, LangChain, LangGraph, CrewAI, or similar.
- Production monitoring experience for AI systems, including behavioral drift detection, output sampling, or shadow scoring.
- Red-teaming or adversarial testing experience for AI models or agents.
- Preferred: experience in identity verification, fraud detection, or regulated industries.
- Preferred: familiarity with Anthropic’s model evaluation methodology or published evaluation research.
- Preferred: experience applying Datadog or OpenTelemetry to AI workloads.
- Preferred: experience building widely adopted developer tooling or platforms.
Benefits
- Comprehensive medical, dental, and vision insurance.
- Health savings account and flexible spending accounts, including medical, limited-purpose, dependent-care, and commuter accounts.
- Basic and voluntary life and AD&D insurance, short- and long-term disability insurance, accident insurance, and critical illness insurance.
- 401(k) with company match.
- Parental leave and unlimited paid time off subject to company policy, including eight company-wide holidays.
- Referral bonus policy, employee assistance program, pet insurance, travel assistance, wellbeing discounts, childcare discounts, benefit advocates, and learning and development benefits.
- Full-time, in-office role requiring five days per week at an ID.me office; this posting is for Mountain View, California.
