
Machine Learning Engineer, RL Environments - New Graduates
Preference Model5 months ago
Toronto, Canada or San Francisco, CA, USAEntry Level
Responsibilities
- Design and build reinforcement-learning environments and reward schemes that produce clean, learnable signals for frontier models.
- Develop expertise across machine-learning research, training infrastructure, and inference infrastructure.
- Collaborate with teammates to create ideas and tools that improve the environment-building process.
Requirements
- Strong machine-learning fundamentals and broad research interests, with the ability to translate research ideas into RLVR problems.
- Proficiency in Python and systems programming; PyTorch or JAX is preferred.
- Strong problem-solving, ownership, responsiveness to feedback, and ability to meet throughput expectations.
- Expertise in an active deep-learning or machine-learning research area, publications, or public code is preferred.
- Research experience, including PhD or MS work, is a plus.
- Deep understanding of transformer internals and experience with kernel development using CUDA, Triton, or Pallas are preferred.
- Research, coursework, or personal projects involving reinforcement-learning environments are preferred.
- Open-source contributions to ML infrastructure or RL tooling are preferred.
- Experience with AWS, GCP, Azure, or infrastructure-as-code tools is preferred.
Benefits
- Competitive cash and equity compensation (specific amounts not stated)
- Health, vision, and dental benefits
- 401(k) match
- Visa sponsorship and relocation support available
- High ownership and autonomy in a fast-moving startup environment
- Opportunity to work with leading machine learning engineers
Categories
AI ResearchML Engineering
About Preference Model
Preference Model builds RL environments that automate ML research and engineering.