about 2 hours ago
Responsibilities
- Build RL and post-training pipelines for multimodal, vision-language, and code-generating agents.
- Develop infrastructure for supervised fine-tuning, preference optimization, reward modeling, and reinforcement learning.
- Create systems for collecting, filtering, replaying, and learning from agent trajectories.
- Design rewards, verifiers, and evaluations for long-horizon agent tasks.
- Improve agents’ ability to plan, write and execute code, use tools, recover from errors, and complete complex workflows.
- Scale distributed training and high-throughput rollout generation across multi-GPU environments.
- Improve training reliability, reproducibility, observability, and cost efficiency.
- Help define Moonlake’s long-term strategy for agent, robotics, and embodied-model training.
Requirements
- Real-world experience training large language, vision-language, multimodal, or code models.
- Strong experience in reinforcement learning, post-training, or large-scale fine-tuning.
- Experience building distributed training or high-throughput inference systems.
- Familiarity with supervised fine-tuning, preference optimization, reward modeling, and agentic RL.
- Experience with code-generation agents, long-horizon evaluation, or tool-using systems.
- Strong Python skills and experience with PyTorch, JAX, or similar frameworks.
- Ability to work across data, models, environments, rewards, evaluation, and infrastructure.
- Strong research judgment and a bias toward building reliable systems.
