ID.me

Staff Software Engineer- AI Agent Evaluations

ID.me
Apply
2 days ago
Mountain View, CA, USAStaff+

Base Salary

$218k - $271k/yr

Responsibilities

  • Own the framework for evaluating, validating, and monitoring AI agents across development, staging, and production.
  • Design and maintain evaluation pipelines for LLM outputs, agent behavior, tool use, and multi-turn interactions.
  • Instrument agentic systems to detect behavioral drift, regressions, latency issues, hallucinations, tool misuse, and policy violations.
  • Lead AI test strategies including red-teaming, golden dataset construction, LLM-as-judge pipelines, and property-based testing.
  • Build internal developer tooling, feedback loops, and workflows that accelerate local testing and evaluation runs.
  • Establish quality standards, patterns, and education for safely developing and deploying AI-powered features.
  • Build and maintain RAG pipelines with strategic data chunking and pre-retrieval data quality gates.
  • Partner with Security, Platform, Product, and AI/ML teams to embed quality gates into agent workflows.
  • Mentor senior and mid-level engineers on AI evaluation, observability, and testing practices.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, or equivalent experience.
  • 8+ years building and operating production software systems.
  • Demonstrated experience evaluating or testing LLM-powered features or autonomous agents in production.
  • Proficiency with AI-assisted development tools such as Claude Code, Cursor, or equivalent.
  • Strong backend engineering fundamentals in Python, Java, Go, or equivalent.
  • Experience designing test infrastructure, quality gates, or evaluation pipelines at scale.
  • Experience improving developer experience through internal tooling, toil reduction, or engineering workflow acceleration.
  • Ability to lead cross-team technical initiatives and influence engineering standards.
  • Strong written and verbal communication across engineering, product, and leadership.
  • Experience building LLM-agent evaluation frameworks, including correctness graders, LLM-as-judge systems, human-in-the-loop evaluations, or benchmark dataset curation.
  • Familiarity with agentic frameworks such as Claude API, Anthropic SDK, BrainTrust, LangChain, LangGraph, CrewAI, or similar.
  • Production monitoring experience for AI systems, including behavioral drift detection, output sampling, or shadow scoring.
  • Red-teaming or adversarial testing experience for AI models or agents.
  • Preferred: experience in identity verification, fraud detection, or regulated industries.
  • Preferred: familiarity with Anthropic’s model evaluation methodology or published evaluation research.
  • Preferred: experience applying Datadog or OpenTelemetry to AI workloads.
  • Preferred: experience building widely adopted developer tooling or platforms.

Benefits

  • Comprehensive medical, dental, and vision insurance.
  • Health savings account and flexible spending accounts, including medical, limited-purpose, dependent-care, and commuter accounts.
  • Basic and voluntary life and AD&D insurance, short- and long-term disability insurance, accident insurance, and critical illness insurance.
  • 401(k) with company match.
  • Parental leave and unlimited paid time off subject to company policy, including eight company-wide holidays.
  • Referral bonus policy, employee assistance program, pet insurance, travel assistance, wellbeing discounts, childcare discounts, benefit advocates, and learning and development benefits.
  • Full-time, in-office role requiring five days per week at an ID.me office; this posting is for Mountain View, California.

Tech Stack

ID.me

About ID.me

1,001-5,000 employees
Contact me