Mercor

Research Engineer – Benchmarking

Mercor
Apply
4 hours ago

Base Salary

$130k - $500k/yr

Responsibilities

  • Design, implement, and maintain benchmarks and metrics for tool use, agentic behavior, and real-world reasoning.
  • Build and operate end-to-end LLM evaluation systems, including runs, scoring, dashboards, and reporting.
  • Conduct systematic failure analysis of model outputs, categorize failure modes, quantify their prevalence, and inform reward design, data curation, and benchmark design.
  • Create and refine rubrics, automated evaluators, and scoring frameworks, balancing rigor, scalability, human evaluation, model-as-judge evaluation, calibration, and agreement.
  • Measure data usability, quality, and impact on key benchmarks and guide data generation, augmentation, and curation.
  • Collaborate with AI researchers, applied AI teams, and data producers to align evaluations with training objectives.
  • Own benchmarking, evaluation, and failure-analysis workflows in a fast-paced research environment.

Requirements

  • Strong applied research background focused on model evaluation, benchmarking, and/or failure analysis.
  • Strong coding skills and hands-on experience with ML models and evaluation code.
  • Solid understanding of data structures, algorithms, and backend systems.
  • Comfort working with APIs, SQL/NoSQL, and cloud platforms for running and storing evaluation results.
  • Ability to reason about model behavior, experimental results, and data quality from evaluations and failure analyses.
  • Industry experience on a post-training or evaluation/benchmarking team is preferred.
  • Publications at top-tier venues such as NeurIPS, ICML, or ACL are preferred, especially in evaluation or benchmarking.
  • Experience building or running LLM evaluations, benchmarks, or failure-analysis pipelines is preferred.
  • Experience with synthetic data generation, rubric design, or RL-style workflows using evaluations for reward shaping is preferred.
  • Work samples or code demonstrating relevant evaluation frameworks, benchmark suites, failure-analysis reports, or tooling are preferred.

Benefits

  • In-person work five days a week in San Francisco, New York City, or London.
  • Bi-annual performance bonus structure.
  • Generous equity grant vested over four years.
  • Up to $15k relocation bonus.
  • $10K housing bonus for employees living within 0.5 miles of the office.
  • $1.5K monthly meal stipend.
  • Free Equinox membership.
  • $200 monthly laundry reimbursement.
  • $200 monthly personal wellness reimbursement.
  • Health, dental, and vision insurance.

Tech Stack

Categories

AI ResearchBackend
Mercor

About Mercor

201-500 employees

We find the best experts in every professional domain and put their knowledge to work training frontier models. Through APEX, we measure whether those models can actually perform economically valuable work. We're also bringing that expertise to enterprises: deploying custom AI agents, staffing teams with vetted domain experts, and helping organizations encode their own knowledge into AI systems.