3 hours ago
Responsibilities
- Design evaluations for voice, vision, and browser use to guide improvements in the agent, context, and models.
- Build and maintain internal benchmarks and simulations.
- Develop systems to score customer-specific scenarios and identify gaps in current capabilities.
- Use evaluation findings to optimize the agent, context, and models.
Requirements
- Engineering experience with statistics and/or applied machine learning.
- Comfort working in a production codebase.
- Prior experience with evaluations, voice, realtime systems, or browser/computer-use agents is a bonus.
