16 days ago
Responsibilities
- Design and build a unified evaluation system for agentic and algorithmic solutions across task success, quality, robustness, motifs, agent topologies, and decomposition strategies.
- Mine real commits and traces and create task suites covering code search, code editing and repair, repository summarization, and tool use.
- Implement sandboxed execution grading and maintain reproducible, contamination-controlled, stable benchmarks and baselines.
- Compare agentic and algorithmic approaches based on resolved-task quality and transparently account for added steps, latency, and cost.
- Lead external benchmark co-publications and ensure results withstand peer and customer review.
- Use evaluation results to inform shipping decisions, customer proof points, and company-wide quality claims.
Requirements
- PhD in computer science, machine learning, or a related field, or an equivalent research track record.
- Authorship or co-authorship of a benchmark or evaluation paper at a recognized venue such as NeurIPS Datasets and Benchmarks, ICML, ICLR, or ACL.
- Working understanding of contamination, benchmark overfitting, weak baselines, underpowered comparisons, and irreproducible results in LLM and agent evaluation.
- Strong Python engineering skills and comfort with sandboxed and distributed execution and CI.
- Hands-on experience building or rigorously evaluating agentic or multi-step LLM systems.
Benefits
- Competitive salary determined by skills and experience.
- Equity and ownership.
- Private healthcare.
- Visa sponsorship and relocation benefits.
- In-person work at the London office with provided tools, workspace, and setup.
