5 hours ago
Base Salary
$125k - $200k/yr
Responsibilities
- Design and build reinforcement learning environments, reward functions, and training pipelines.
- Train and fine-tune models using PPO, GRPO, DPO, RLHF, and RLAIF.
- Develop evaluation frameworks for model and agent performance.
- Run experiments, interpret results, and determine which approaches to pursue.
- Scale training on GPU clusters and maintain reliable pipelines.
- Translate research ideas into production systems.
- Help establish engineering culture and hire future engineers.
Requirements
- At least 2 years of hands-on reinforcement learning or machine learning engineering experience, with relevant experience potentially ranging from 2 to 10 or more years.
- Strong Python skills and deep experience with PyTorch or JAX.
- Practical experience training models with reinforcement learning, including policy gradient methods, reward modeling, or RLHF.
- Experience with distributed training and GPU infrastructure, including Ray, CUDA, and Kubernetes.
- Experience with LLM post-training or agent training is preferred.
- Familiarity with Gymnasium, Ray RLlib, Isaac, MuJoCo, TRL, verl, or OpenRLHF is preferred.
- A degree in computer science, mathematics, physics, or a related field is sought; a master's or PhD is a plus.
- Publications or open-source work in reinforcement learning are valued.
- Ability to work hands-on and move quickly in a small, early-stage team.
Benefits
- Compensation is $125,000 to $200,000 USD annually.
- On-site role in San Francisco, California, United States.
- Opportunity to work directly with the founders as a founding engineer on an early-stage AI team.
