Clera

Research Engineer, Benchmarks

Clera
Apply
1 day ago
Singapore, SingaporeMid Level / Senior

Base Salary

$150k - $250k/yr

Responsibilities

  • Design, implement, and maintain internal benchmarks for evaluating frontier AI agents on domain-specific tasks.
  • Partner with subject-matter experts to define realistic workflows and translate them into evaluation tasks.
  • Build infrastructure to run models and agents against benchmark tasks at scale.
  • Develop metrics and statistical analyses measuring benchmark difficulty, reliability, and failure modes.
  • Validate benchmark performance against real-world evaluations and customer expectations.
  • Write technical documentation and benchmark reports for research and engineering audiences.

Requirements

  • 2 to 4 years of experience in research engineering or machine learning engineering focused on AI benchmarks, evaluation infrastructure, or agent environments.
  • Strong proficiency in Python, Docker, and Linux.
  • Hands-on experience designing and running benchmarks or evaluation environments for AI agents or large language models.
  • Experience developing metrics and validation studies for benchmark difficulty, reliability, and real-world correlation.
  • Experience collaborating with domain experts to translate workflows into evaluation criteria.
  • Strong technical writing skills and the ability to reason about task design, scoring, and edge cases.
  • Experience with reinforcement learning training pipelines, data generation, or reinforcement-learning agent evaluation is a plus.
  • Published papers or blog posts on AI benchmarking, model evaluation, or failure modes are a plus.
  • Background at a frontier AI lab, research institution, or widely used public benchmark project is a plus.
  • Ability to work independently in a fast-paced, early-stage environment with unstructured problem spaces.

Benefits

  • Salary range of $150,000 to $250,000 USD annually.
  • Visa sponsorship is available.
  • On-site work in Singapore.
Contact me