2 hours ago
Remote, United States +2 moreMid Level
H1B sponsor
Base Salary
$200k - $350k/yr
Responsibilities
- Design post-training systems and methodologies for frontier models, including supervised fine-tuning, reinforcement learning, preference optimization, and reward modeling.
- Translate research and partner needs into hypotheses, experiments, evaluation plans, and production-quality implementations.
- Build evaluation frameworks, benchmarks, training environments, data-processing pipelines, and quality-control systems.
- Run iterative experimentation, evaluate and interpret results, and apply learnings to systems and products.
- Partner with AI researchers and domain experts to develop high-signal data, feedback, and evaluation methods.
- Productize repeatable patterns into reusable software and platforms.
- Contribute through technical design, code quality, mentorship, research, benchmarks, open-source tools, and technical writing.
Requirements
- At least 3 years of demonstrated strength in post-training, fine-tuning, or model-evaluation work.
- Relevant experience may include reinforcement learning, supervised fine-tuning, LoRA/PEFT, full fine-tuning, RLHF, DPO, PPO, reward modeling, or training environments.
- Strong Python skills and the ability to write clean, efficient, scalable software.
- Hands-on experience with modern machine-learning tooling, particularly PyTorch and large-scale data, training, or evaluation workflows.
- Strong experimental judgment, including hypothesis formation, metric selection, failure diagnosis, and signal-versus-noise assessment.
- Experience designing systems and making tradeoffs involving quality, scale, reliability, and reuse.
- Ability to work with substantial ownership in an ambiguous, fast-moving environment.
- Collaborative communication skills and the ability to work with researchers, engineers, domain experts, and customers.
- Preferred experience includes large-scale ML training, inference, data, or evaluation systems; LLM or agent benchmarks; annotation systems; data-quality frameworks; reinforcement learning; alignment; model behavior; synthetic data; human-in-the-loop systems; published research; open-source contributions; technical leadership; or productizing research into reusable platforms.
Benefits
- Full-time US employees receive equity, 401(k) matching, competitive compensation, financial coaching, paid parental leave, fertility benefits, parental coaching, medical, dental and vision coverage, mental health support, and a $500 wellness stipend.
- Benefits include a $2,000 learning stipend, ongoing development, commuting support, free lunch and gym access at the San Francisco office, flexible PTO, 15 holidays, 2 flex days, team outings, and referral bonuses.
- San Francisco and Mountain View are preferred locations, with exceptional candidates in other locations considered.
Categories
AI ResearchML Engineering
About Handshake
Handshake builds a career network and recruiting platform that connects college students and recent grads with employers, sold as SaaS to universities and subscriptions/solutions to employers. Founded in 2014 and headquartered in San Francisco, it serves 1,600+ educational institutions and over 1 million employers. The company also operates Handshake AI, which partners with frontier AI labs on human data collection and evaluations for model training and post-training workflows.
