4 hours ago
Responsibilities
- Build new reinforcement-learning environments for different agent capabilities and industry areas.
- Train and evaluate AI agents within reinforcement-learning environments.
- Integrate tasks, data, tool implementations, and verifiers.
- Collaborate across modeling and product teams to identify agent-performance gaps and improve agents and environments.
- Work with external vendors to create expert-built reinforcement-learning environments.
- Build tools to ensure task, data, and verifier quality.
- Automate discovery of model capability gaps and systematically measure agent performance during evaluations and training.
Requirements
- Experience engineering agents and optimizing them for specific industry use cases.
- Extensive hands-on experience reviewing agent trajectories to identify failure points and resolving them through model training or harness engineering.
- Experience measuring agent capabilities and creating repeatable evaluation processes.
- Experience defining desired agent outcomes, implementing verifiers, and tuning reward designs.
- Experience designing and running annotation workflows to analyze agent performance and verify data quality.
- Experience building synthetic data pipelines to scale evaluation and training efforts.
- Regular use of agents with demonstrated improvements to personal workflows.
- Experience with reinforcement-learning training, including scaling, troubleshooting, and tuning environments, is a plus.
Benefits
- Weekly lunch stipend of $75/£75 or equivalent in local currency.
- Full health and dental benefits, including a separate mental health budget.
- RRSP matching, 401K, and Pension Scheme.
- 100% parental leave top-up for up to 6 months for either parent.
- Annual enrichment benefits covering arts and culture, fitness and wellness, quality time, and workspace improvement.
- Education and learning stipend for conferences, courses, and coaching.
- Six weeks of paid vacation, or 30 working days.
- Travel budget for remote employees to visit other offices and an annual company offsite.
- Remote-friendly work arrangement with offices in Toronto, London, New York City, San Francisco, Montreal, Paris, Berlin, and Seoul.
- Co-working benefit for employees not near an office.
- Daily lunch program, snacks, and community and social events for office-based employees.
- $500 home office stipend.
Categories
AI Research
About Cohere
Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation models and end-to-end AI products designed to solve real-world business problems. We partner closely with companies to deliver seamless integration, full customization, and easy-to-use solutions for their workforce and customers. Our all-in-one platform offers enterprises the highest levels of data security, privacy and optionality to deploy across all major cloud providers, private cloud environments, or on-premises. HQ: 171 John Street, 2nd Floor, Toronto, ON M5T 1X3
