
Machine Learning Engineer, RL Environments
Preference Model5 months ago
Toronto, Canada or San Francisco, CA, USASenior
Responsibilities
- Design and build reinforcement learning environments for ML research and engineering tasks.
- Develop robust reward functions that produce clean, learnable signals for frontier models.
- Build expertise across machine learning research, training, and inference infrastructure.
- Collaborate on new ideas and tools to improve the environment-building process.
Requirements
- Strong machine learning fundamentals and broad research interests, with the ability to translate research concepts into RLVR problems.
- Proficiency in Python and systems programming, plus proficiency in at least one of PyTorch or JAX.
- Ability to solve problems, take ownership, and drive solutions end-to-end.
- Passion for staying current with the evolving machine learning infrastructure landscape.
- Ability to meet throughput expectations and respond quickly to feedback.
- Preferred qualifications include expertise in an active deep learning or machine learning research area with publications or public code.
- Research experience, including a PhD or MS, is a significant plus.
- Preferred experience includes transformer internals, modern LLM training and inference, inference libraries such as vLLM or SGLang, kernel development with CUDA, Triton, or Pallas, or building complex interactive reinforcement learning environments.
Benefits
- Health, vision, and dental benefits
- 401(k) match
- Visa sponsorship and relocation support available
- Ownership and autonomy in a fast-moving startup environment
- Opportunity to work with top machine learning engineers
Categories
About Preference Model
Preference Model builds RL environments that automate ML research and engineering.