Hark

Member of Technical Staff, Mid-training

Hark
Apply
10 days ago
San Jose, CA, USASenior

Base Salary

$180k - $450k/yr

Responsibilities

  • Design and implement mid-training strategies for reasoning, planning, tool use, and long-horizon decision-making.
  • Scale synthetic data generation pipelines for coding, agent trajectories, and multimodal data, and optimize data mixtures for downstream reinforcement learning.
  • Build and optimize distributed training pipelines for large models across GPU clusters.
  • Develop evaluation frameworks for task success, reasoning quality, and tool-use accuracy.
  • Conduct experiments and ablations to study training dynamics, scaling behavior, and bottlenecks.
  • Collaborate with pre-training, post-training, and product teams on real-world agent use cases.
  • Drive innovation in long-context learning, data distillation, and training efficiency while contributing to the model roadmap.

Requirements

  • Strong machine learning background with hands-on experience training or fine-tuning large language, multimodal, or equivalent models.
  • Deep understanding of reinforcement learning, including policy optimization, reward design, exploration, and environment design.
  • Experience with simulation or execution environments such as code interpreters, sandboxed execution, game environments, or robotics simulators.
  • Ability to design rigorous experiments and diagnose training failures and scaling bottlenecks.
  • Proficiency in Python and PyTorch, with comfort working across research and systems code.
  • Ability to work effectively in a fast-moving, research-forward environment.
  • Relevant backgrounds may include RL research, robotics, competitive programming systems, compilers, formal methods, or large-scale machine learning.
  • Bonus qualifications include mid-training or post-training experience, synthetic data and trajectory-based training, training models with 100B+ parameters, agent frameworks, open-source ML systems, top-conference publications, and distributed training optimization.

Benefits

  • Full-time position with a US base salary range of $180,000-$450,000 annually.
  • Total compensation may include additional components and benefits depending on the role.

Tech Stack

Categories

AI ResearchML Engineering
Hark

About Hark

1-10 employees
Contact me