
Preference Model
Preference Model builds RL environments that automate ML research and engineering.
Open Positions at Preference Model
4 open positions
Build reinforcement learning environments that teach frontier language models to reason about and solve real-world cybersecurity problems. You’ll combine hands-on security expertise with software engineering to create verifiable environments for vulnerability discovery, exploitation, patching, and reverse engineering.
Build reinforcement learning environments and reward functions that help frontier LLMs develop stronger reasoning and machine learning capabilities. This high-ownership role combines ML engineering, research, and infrastructure work on real-world training tasks.
Build reinforcement-learning environments and reward systems that help frontier models learn to reason on complex ML research and engineering tasks. This entry-level role offers new and recent graduates direct ownership in an ML research engineering startup.
Build sophisticated reinforcement-learning environments that reveal and improve frontier models’ ability to handle complex software-engineering work. This high-ownership role combines deep systems engineering, coding-agent evaluation, and end-to-end task design.