Clera

Research Engineer, Benchmarks

Clera
Apply
3 days ago
Singapore, SingaporeMid Level / Senior

Base Salary

$150k - $250k/yr

Responsibilities

  • Design, implement, and own internal benchmarks for evaluating frontier AI agents on domain-specific tasks.
  • Partner with subject-matter experts to translate realistic workflows into benchmark tasks and evaluation criteria.
  • Build and operate infrastructure for running models and agents against benchmark tasks at scale.
  • Develop metrics and statistical analyses for benchmark difficulty, reliability, and failure modes.
  • Validate benchmark performance against real-world evaluations, customer needs, and frontier lab expectations.
  • Produce technical documentation and benchmark reports for research and engineering audiences.

Requirements

  • 2–4 years of experience in software engineering, ML engineering, or research roles with a focused track record in AI benchmarks or evaluation infrastructure.
  • Strong proficiency in Python, Docker, and Linux environments.
  • Experience designing, implementing, and running benchmarks or evaluation environments for AI agents or large language models.
  • Experience building infrastructure to reliably run AI models or agents against benchmark or evaluation tasks.
  • Ability to analyze and model workflows across diverse technical or business domains.
  • Strong attention to detail and ability to reason from first principles about task design, scoring, and failure modes.
  • Strong written communication skills and experience producing technical documentation or benchmark reports.
  • Ability to work effectively in unstructured problem spaces at an early-stage startup.
  • Experience with reinforcement learning pipelines, data generation, or RL agent evaluation is a bonus.
  • Published work on AI benchmarking or model evaluation is a bonus.

Benefits

  • Annual salary range of USD 150,000 to 250,000.
  • Visa sponsorship is available.
  • On-site role in Singapore.

Categories

Contact me