Preference Model

Machine Learning Engineer, RL Environments - New Graduates

Preference Model
Apply
5 months ago
Toronto, Canada or San Francisco, CA, USAEntry Level

Responsibilities

  • Design and build reinforcement-learning environments and reward schemes that produce clean, learnable signals for frontier models.
  • Develop expertise across machine-learning research, training infrastructure, and inference infrastructure.
  • Collaborate with teammates to create ideas and tools that improve the environment-building process.

Requirements

  • Strong machine-learning fundamentals and broad research interests, with the ability to translate research ideas into RLVR problems.
  • Proficiency in Python and systems programming; PyTorch or JAX is preferred.
  • Strong problem-solving, ownership, responsiveness to feedback, and ability to meet throughput expectations.
  • Expertise in an active deep-learning or machine-learning research area, publications, or public code is preferred.
  • Research experience, including PhD or MS work, is a plus.
  • Deep understanding of transformer internals and experience with kernel development using CUDA, Triton, or Pallas are preferred.
  • Research, coursework, or personal projects involving reinforcement-learning environments are preferred.
  • Open-source contributions to ML infrastructure or RL tooling are preferred.
  • Experience with AWS, GCP, Azure, or infrastructure-as-code tools is preferred.

Benefits

  • Competitive cash and equity compensation (specific amounts not stated)
  • Health, vision, and dental benefits
  • 401(k) match
  • Visa sponsorship and relocation support available
  • High ownership and autonomy in a fast-moving startup environment
  • Opportunity to work with leading machine learning engineers
Preference Model

About Preference Model

11-50 employees

Preference Model builds RL environments that automate ML research and engineering.

Contact me