2 months ago
Base Salary
$130k - $500k/yr
Responsibilities
- Design, implement, and maintain benchmarks and metrics for tool use, agentic behavior, and real-world reasoning.
- Build and operate end-to-end LLM evaluation systems, including runs, scoring, dashboards, and reporting.
- Conduct systematic failure analysis of model outputs, categorize failure modes, quantify their prevalence, and inform reward design, data curation, and benchmark design.
- Create and refine rubrics, automated evaluators, and scoring frameworks, balancing rigor, scalability, human evaluation, model-as-judge evaluation, calibration, and agreement.
- Measure data usability, quality, and impact on key benchmarks and guide data generation, augmentation, and curation.
- Collaborate with AI researchers, applied AI teams, and data producers to align evaluations with training objectives.
- Own benchmarking, evaluation, and failure-analysis workflows in a fast-paced research environment.
Requirements
- Strong applied research background focused on model evaluation, benchmarking, and/or failure analysis.
- Strong coding skills and hands-on experience with ML models and evaluation code.
- Solid understanding of data structures, algorithms, and backend systems.
- Comfort working with APIs, SQL/NoSQL, and cloud platforms for running and storing evaluation results.
- Ability to reason about model behavior, experimental results, and data quality from evaluations and failure analyses.
- Industry experience on a post-training or evaluation/benchmarking team is preferred.
- Publications at top-tier venues such as NeurIPS, ICML, or ACL are preferred, especially in evaluation or benchmarking.
- Experience building or running LLM evaluations, benchmarks, or failure-analysis pipelines is preferred.
- Experience with synthetic data generation, rubric design, or RL-style workflows using evaluations for reward shaping is preferred.
- Work samples or code demonstrating relevant evaluation frameworks, benchmark suites, failure-analysis reports, or tooling are preferred.
Benefits
- In-person work five days a week in San Francisco, New York City, or London.
- Bi-annual performance bonus structure.
- Generous equity grant vested over four years.
- Up to $15k relocation bonus.
- $10K housing bonus for employees living within 0.5 miles of the office.
- $1.5K monthly meal stipend.
- Free Equinox membership.
- $200 monthly laundry reimbursement.
- $200 monthly personal wellness reimbursement.
- Health, dental, and vision insurance.
About Mercor
Mercor builds an expert-powered platform that trains, evaluates, and deploys AI systems for AI labs and enterprises. It operates APEX to assess model performance on economically valuable tasks, and provides services such as custom AI agents and teams of vetted domain experts who encode organizational knowledge into AI. Founded in 2023 and headquartered in San Francisco, it is privately held and works with frontier AI labs and enterprise clients requiring strict data isolation.
