4 hours ago
Singapore, SingaporeMid Level / Senior
Base Salary
$150k - $250k/yr
Responsibilities
- Design, implement, and own internal benchmarks for evaluating frontier agents on domain-specific tasks.
- Partner with subject-matter experts to define realistic workflows and evaluation criteria.
- Build scalable infrastructure to run models and agents against benchmark tasks using Python, Docker, and Linux.
- Develop metrics and statistical analyses for benchmark difficulty, reliability, and failure modes.
- Validate benchmark performance against real-world evaluations and customer needs.
- Write technical documentation and benchmark reports for research and engineering audiences.
Requirements
- 2 to 4 years of experience in research engineering or machine learning engineering, focused on AI benchmarks, evaluation infrastructure, or agent environments.
- Strong proficiency in Python, Docker, and Linux.
- Experience designing and running benchmarks or evaluation environments for AI agents or large language models.
- Experience developing metrics, statistical analyses, or validation studies for benchmark quality and real-world correlation.
- Experience translating workflows into structured evaluation tasks with domain experts.
- Strong technical writing skills; published papers or technical posts on AI benchmarking, model evaluation, or failure modes are a plus.
- Ability to reason from first principles about task design, scoring, and edge cases.
- Ability to work independently in fast-paced, early-stage startup environments.
- Experience with reinforcement learning training pipelines, data generation, or RL agent evaluation is a bonus.
Benefits
- Salary range of $150,000 to $250,000 USD annually.
- Visa sponsorship is available.
- Full-time, in-person role on-site in Singapore.
