5 months ago
San Francisco, CA, USA or New York, NY, USASenior
Responsibilities
- Design and build AI evaluation systems with curated datasets, offline replay, scorers or judges, regression alerts, and dashboards.
- Develop feedback loops by collecting, cleaning, and interpreting user signals to guide model and harness changes.
- Build analysis tooling for debugging agent behavior, investigating failure modes, clustering themes, and surfacing actionable insights.
- Define and operationalize agent quality states, alerting, reliability guardrails, and triage primitives.
- Partner with research, product, data, and infrastructure teams to improve model quality, cost, and agent behavior.
Requirements
- Experience building and operating evaluation or measurement systems such as AI evaluations, experimentation, ranking or relevance systems, or search-quality systems.
- Ability to turn ambiguous quality questions into concrete metrics, pipelines, and decisions.
- Strong data acumen and ability to collaborate with data scientists and researchers.
- Strong software engineering fundamentals and enthusiasm for shipping production systems.
- Interest in model and agent behavior and awareness of emerging research and industry trends.
