10 days ago
San Jose, CA, USASenior
Base Salary
$180k - $450k/yr
Responsibilities
- Design and build RL environments and tasks for long-horizon planning, tool and computer use, multimodal interaction, and hardware-based agent behavior.
- Architect reward functions, graders, and task curricula that determine what environments teach or reveal.
- Scale environment infrastructure to run thousands of tasks and rollouts in parallel across simulated settings and real agentic or hardware sessions.
- Influence major training-run decisions and validate whether resulting model improvements work in practice.
- Build self-improvement loops in which models generate, grade, and refine their own training environments.
- Translate ambiguous behavioral concerns into hypotheses, experiments, pipelines, analyses, and follow-up decisions.
Requirements
- Strong machine-learning background with hands-on experience training or fine-tuning large language, multimodal, or equivalent models.
- Hands-on experience building RL environments, simulators, or task suites such as Gym-style interfaces, game engines, browser or OS automation sandboxes, or robotics and hardware simulators.
- Direct experience with LLMs and post-training techniques including RL, RLHF/RLAIF, reward modeling, graders, synthetic data pipelines, or agentic and tool-using systems.
- Track record solving ambiguous problems involving loosely defined objectives and noisy data through hands-on engineering and judgment.
- Ability to form a strong point of view on useful, trustworthy, and pleasant personal-assistant behavior.
- Ability to collaborate across research, product, infrastructure, data, hardware, and safety teams and communicate clearly between them.
- Preferred experience with RLHF, DPO, GRPO, PPO, or similar algorithms applied to language, code, or agentic settings.
- Preferred familiarity with agent benchmarks and evaluation environments including OSWorld, Toolathlon, and GDPval.
- Preferred research or publication-level experience in reward-model design, reward signals, or reward-hacking mitigation.
- Preferred experience with trajectory-based training, imitation learning, data distillation, computer use, GUI agents, multimodal tool-using systems, or 100B+ parameter model training.
- Open-source ML contributions or publications at venues such as NeurIPS, ICML, ICLR, EMNLP, or COLM are preferred.
Benefits
- Full-time position.
- US base salary range of $180,000-$450,000 annually.
- Total compensation may include additional components and benefits depending on the role.
