GrepJob
Moonlake

Member of Technical Staff — Agent Post-Training

Moonlake
Apply
about 2 hours ago
San Francisco, CA, USAMid Level / Senior
H1B Sponsor

Responsibilities

  • Build RL and post-training pipelines for multimodal, vision-language, and code-generating agents.
  • Develop infrastructure for supervised fine-tuning, preference optimization, reward modeling, and reinforcement learning.
  • Create systems for collecting, filtering, replaying, and learning from agent trajectories.
  • Design rewards, verifiers, and evaluations for long-horizon agent tasks.
  • Improve agents’ ability to plan, write and execute code, use tools, recover from errors, and complete complex workflows.
  • Scale distributed training and high-throughput rollout generation across multi-GPU environments.
  • Improve training reliability, reproducibility, observability, and cost efficiency.
  • Help define Moonlake’s long-term strategy for agent, robotics, and embodied-model training.

Requirements

  • Real-world experience training large language, vision-language, multimodal, or code models.
  • Strong experience in reinforcement learning, post-training, or large-scale fine-tuning.
  • Experience building distributed training or high-throughput inference systems.
  • Familiarity with supervised fine-tuning, preference optimization, reward modeling, and agentic RL.
  • Experience with code-generation agents, long-horizon evaluation, or tool-using systems.
  • Strong Python skills and experience with PyTorch, JAX, or similar frameworks.
  • Ability to work across data, models, environments, rewards, evaluation, and infrastructure.
  • Strong research judgment and a bias toward building reliable systems.

Tech Stack

Categories

AI & MLBackendData Science