1 day ago
Singapore, SingaporeMid Level / Senior
Base Salary
$150k - $250k/yr
Responsibilities
- Design, implement, and maintain internal benchmarks for evaluating frontier AI agents on domain-specific tasks.
- Partner with subject-matter experts to define realistic workflows and translate them into evaluation tasks.
- Build infrastructure to run models and agents against benchmark tasks at scale.
- Develop metrics and statistical analyses measuring benchmark difficulty, reliability, and failure modes.
- Validate benchmark performance against real-world evaluations and customer expectations.
- Write technical documentation and benchmark reports for research and engineering audiences.
Requirements
- 2 to 4 years of experience in research engineering or machine learning engineering focused on AI benchmarks, evaluation infrastructure, or agent environments.
- Strong proficiency in Python, Docker, and Linux.
- Hands-on experience designing and running benchmarks or evaluation environments for AI agents or large language models.
- Experience developing metrics and validation studies for benchmark difficulty, reliability, and real-world correlation.
- Experience collaborating with domain experts to translate workflows into evaluation criteria.
- Strong technical writing skills and the ability to reason about task design, scoring, and edge cases.
- Experience with reinforcement learning training pipelines, data generation, or reinforcement-learning agent evaluation is a plus.
- Published papers or blog posts on AI benchmarking, model evaluation, or failure modes are a plus.
- Background at a frontier AI lab, research institution, or widely used public benchmark project is a plus.
- Ability to work independently in a fast-paced, early-stage environment with unstructured problem spaces.
Benefits
- Salary range of $150,000 to $250,000 USD annually.
- Visa sponsorship is available.
- On-site work in Singapore.
