about 1 month ago
Base Salary
$250k - $350k/yr
Responsibilities
- Partner with AI lab researchers to translate post-training goals and ambiguous research questions into scoped, executable projects.
- Design and deliver evaluation frameworks, annotation pipelines, and benchmark infrastructure tailored to partner training methodologies.
- Prototype experiments, run evaluations, and interpret results through rapid feedback loops with research partners.
- Make scalable design decisions about data quality and evaluation methodology.
- Mentor engineers and researchers, establish technical standards, and document repeatable patterns for future deployments.
- Stay current on reinforcement learning, post-training, and benchmarking developments and apply relevant insights to customer engagements.
Requirements
- At least 6 years of experience in applied ML, AI research engineering, or a closely related field, including exposure to model training workflows and post-training techniques.
- Strong Python skills and experience across data processing, model evaluation, experiment tracking, and pipeline tooling.
- Working knowledge of reinforcement learning and post-training concepts including RLHF, DPO, and PPO.
- Hands-on experience fine-tuning or lightly optimizing ML models using tools or methods such as Tinker, LoRA, or PEFT.
- Experience with ML data pipelines, data-labeling systems, evaluation frameworks, and quality metrics.
- Strong communication, stakeholder management, prioritization, and technical project leadership skills in ambiguous, fast-moving environments.
- Preferred experience includes LLM evaluation design, production RLHF pipelines, published research or benchmarking, open-source AI/ML tooling, forward-deployed or solutions engineering work, and annotation or human-feedback tooling at scale.
Benefits
- Full-time US employees receive equity, 401(k) matching, medical, dental, and vision coverage, mental health support, paid parental leave, fertility benefits, parental coaching, a $500 wellness stipend, a $2,000 learning stipend, commuting support, free lunch, gym access, flexible PTO, 15 holidays plus 2 flex days, team outings, and referral bonuses.
- The role is based in San Francisco, California, with a hybrid schedule requiring three days per week in the office.