Preference Model

Preference Model

Visit websiteLinkedIn11-50 employees

Preference Model builds RL environments that automate ML research and engineering.

Open Positions at Preference Model

4 open positions

Build reinforcement learning environments that teach frontier language models to reason about and solve real-world cybersecurity problems. You’ll combine hands-on security expertise with software engineering to create verifiable environments for vulnerability discovery, exploitation, patching, and reverse engineering.

4 months ago
Toronto, Canada or San Francisco, CA, USASenior

Build reinforcement learning environments and reward functions that help frontier LLMs develop stronger reasoning and machine learning capabilities. This high-ownership role combines ML engineering, research, and infrastructure work on real-world training tasks.

5 months ago
Toronto, Canada or San Francisco, CA, USAEntry Level

Build reinforcement-learning environments and reward systems that help frontier models learn to reason on complex ML research and engineering tasks. This entry-level role offers new and recent graduates direct ownership in an ML research engineering startup.

Build sophisticated reinforcement-learning environments that reveal and improve frontier models’ ability to handle complex software-engineering work. This high-ownership role combines deep systems engineering, coding-agent evaluation, and end-to-end task design.

5 months ago
Contact me