5 hours ago
Base Salary
$220k - $425k/yr
Responsibilities
- Define golden sets by decomposing real tasks and encoding expert quality standards.
- Build calibrated, difficult-to-game verifiers for agent trajectories and outputs.
- Build and scale the evaluation platform for offline environments, task suites, and grading.
- Analyze production trajectories and turn failure modes into regression tests.
- Run optimization loops across models, prompts, skills, and harnesses.
- Own rollout gates that determine whether agent changes ship.
- Partner with the Enterprise Platform team and Applied AI engineers embedded with customers.
Requirements
- Professional, academic, or research experience in agent engineering and evaluation, including understanding of agent runtimes, harnesses, trajectories, and failure modes.
- Experience building evaluation suites for LLM or agent systems.
- Familiarity with constructing and identifying weaknesses in terminal-bench, tau-bench, and APEX benchmarks.
- Strong judgment in task and rubric design, translating quality standards into measurable evaluations.
- Strong software engineering fundamentals and the ability to work independently on ambiguous, loosely specified problems.
- Experience with Harbor environments and RL environments is preferred.
Benefits
- In-person work five days per week in San Francisco, New York City, or London offices.
- Free Equinox membership.
- Health, dental, and vision insurance.
Categories
About Mercor
We find the best experts in every professional domain and put their knowledge to work training frontier models. Through APEX, we measure whether those models can actually perform economically valuable work. We're also bringing that expertise to enterprises: deploying custom AI agents, staffing teams with vetted domain experts, and helping organizations encode their own knowledge into AI systems.
