3 hours ago
Singapore, SingaporeMid Level
Base Salary
$150k - $250k/yr
Responsibilities
- Design, implement, and maintain benchmarks for evaluating AI agents on domain-specific tasks.
- Collaborate with subject-matter experts to translate real workflows into benchmark tasks and evaluation criteria.
- Build and operate infrastructure for running models and agents against tasks at scale.
- Develop metrics and analyses to assess benchmark difficulty, reliability, and failure modes.
- Validate benchmark results against real-world performance and technical user needs.
- Write documentation and reports for research and engineering audiences.
Requirements
- Two to four years of experience in research engineering, machine learning engineering, or a related technical role, including at least two years building AI benchmarks, evaluation infrastructure, or agent environments.
- Strong Python skills and practical experience with Docker and Linux.
- Experience designing, implementing, and running benchmarks or evaluation environments for AI agents or large language models.
- Experience collaborating with subject-matter experts and analyzing workflows across technical or business domains.
- Experience developing metrics, statistical analyses, or validation studies for evaluations.
- Strong technical writing, attention to detail, and ability to work independently in an early-stage environment.
- A technical educational background.
- Experience with reinforcement learning pipelines, published evaluation work, or widely used benchmarks is preferred.
Benefits
- Salary range of USD 150,000 to 250,000 annually.
- Visa sponsorship is available.
- On-site role in Singapore, Singapore.
