2 months ago
Responsibilities
- Build reinforcement-learning and post-training pipelines for multimodal, vision-language, and code-generating agents.
- Develop infrastructure for supervised fine-tuning, preference optimization, reward modeling, and reinforcement learning.
- Create systems for collecting, filtering, replaying, and learning from agent trajectories.
- Design rewards, verifiers, and evaluations for long-horizon agent tasks.
- Improve agents’ planning, code writing and execution, tool use, error recovery, and complex workflow completion.
- Scale distributed training and high-throughput rollout generation across multi-GPU environments.
- Improve training reliability, reproducibility, observability, and cost efficiency.
- Help define Moonlake’s long-term strategy for agent, robotics, and embodied-model training.
Requirements
- Real-world experience training large language, vision-language, multimodal, or code models.
- Strong experience in reinforcement learning, post-training, or large-scale fine-tuning.
- Experience building distributed training or high-throughput inference systems.
- Familiarity with supervised fine-tuning, preference optimization, reward modeling, and agentic reinforcement learning.
- Experience with code-generation agents, long-horizon evaluation, or tool-using systems.
- Strong Python skills and experience with PyTorch, JAX, or similar frameworks.
- Ability to work across data, models, environments, rewards, evaluation, and infrastructure.
- Strong research judgment and a bias toward building reliable systems.
- Preferred: experience at a frontier AI lab or organization operating large-scale training systems.
- Preferred: experience with code-model post-training, autonomous coding agents, multimodal models, robotics, simulation, embodied AI, verifiable rewards, outcome-based training systems, or large GPU clusters.
Benefits
- On-site, in-person work in San Francisco
Categories
AI ResearchML Engineering
About Moonlake
Moonlake builds interactive world models and simulation infrastructure that generate, simulate, and reason over 3D environments for embodied AI, robotics, and gaming. Its platform and APIs let researchers and developers create assets, scenes, and digital twins at scale, and interact with them via natural language and multimodal inputs. The company is privately held, founded in 2025, headquartered in San Francisco, and raised a $28M seed round in 2025.
