8 hours ago
Base Salary
$150k - $180k/yr
Responsibilities
- Lead development of systems for evaluating tasks across reinforcement learning environments, synthetic data, benchmarks, and domain-specific workflows.
- Define quality standards, metrics, experiments, and processes for assessing AI-agent outputs.
- Develop scalable synthetic-data validation methods, including failure analysis, task mutation checks, and trajectory audits.
- Partner with research engineers, domain experts, and data vendors to diagnose issues and improve data-generation workflows.
- Build production tools, dashboards, validation pipelines, and feedback loops from research insights.
- Mentor research engineers and promote technical rigor and clear communication.
Requirements
- At least 5 years of research or engineering experience building AI or machine-learning data evaluation and quality systems.
- Experience leading technical teams or projects from problem definition through implementation and iteration.
- Advanced proficiency in Python, Docker, and Linux, with a technical education or background.
- Strong understanding of AI evaluations, post-training, and the qualities of realistic, learnable, diverse, and reliable agent-training data.
- Experience building evaluation infrastructure, benchmarks, synthetic-data pipelines, or validation workflows.
- Ability to translate domain expertise and research insights into scalable review systems and production data pipelines.
- Strong written communication and ability to work independently in an early-stage startup environment.
Benefits
- Salary range of $150,000 to $180,000 annually.
- Visa sponsorship is available.
- On-site role in San Francisco, California, United States.
