Callosum

Research Engineer, Benchmarking - Member of Technical Staff

Callosum
Apply
16 days ago
London, United KingdomSenior
H1B sponsor

Responsibilities

  • Design and build a unified evaluation system for agentic and algorithmic solutions across task success, quality, robustness, motifs, agent topologies, and decomposition strategies.
  • Mine real commits and traces and create task suites covering code search, code editing and repair, repository summarization, and tool use.
  • Implement sandboxed execution grading and maintain reproducible, contamination-controlled, stable benchmarks and baselines.
  • Compare agentic and algorithmic approaches based on resolved-task quality and transparently account for added steps, latency, and cost.
  • Lead external benchmark co-publications and ensure results withstand peer and customer review.
  • Use evaluation results to inform shipping decisions, customer proof points, and company-wide quality claims.

Requirements

  • PhD in computer science, machine learning, or a related field, or an equivalent research track record.
  • Authorship or co-authorship of a benchmark or evaluation paper at a recognized venue such as NeurIPS Datasets and Benchmarks, ICML, ICLR, or ACL.
  • Working understanding of contamination, benchmark overfitting, weak baselines, underpowered comparisons, and irreproducible results in LLM and agent evaluation.
  • Strong Python engineering skills and comfort with sandboxed and distributed execution and CI.
  • Hands-on experience building or rigorously evaluating agentic or multi-step LLM systems.

Benefits

  • Competitive salary determined by skills and experience.
  • Equity and ownership.
  • Private healthcare.
  • Visa sponsorship and relocation benefits.
  • In-person work at the London office with provided tools, workspace, and setup.

Tech Stack

Categories

Callosum

About Callosum

11-50 employees
Contact me