NewsBreak

Machine Learning Engineer, LLM Post-Training

NewsBreak
Apply
3 months ago
Mountain View, CA, USASenior
H1B sponsor

Base Salary

$150k - $230k/yr

Responsibilities

  • Lead LLM post-training across continuous pre-training, supervised fine-tuning, and reinforcement learning, with RL as the primary focus.
  • Design and curate instruction datasets, preference pairs, reward signals, on-policy rollouts, and rejection-sampled completions.
  • Partner with product and business stakeholders to translate use cases into training objectives and plans.
  • Run large-scale training on GPU clusters using data parallelism, FSDP, and relevant tensor or pipeline parallelism techniques.
  • Build and maintain evaluation, reward, and verifier pipelines to measure quality, prevent regressions, and ensure training-serving consistency.
  • Turn promising post-training research techniques into working, production-ready code.

Requirements

  • Hands-on experience personally running continuous pre-training, supervised fine-tuning, and reinforcement-learning training for LLMs, including practical experience with RLHF, PPO, GRPO, DPO, or similar methods.
  • Ability to independently design ML data-preparation strategies covering sourcing, cleaning, filtering, labeling, and synthetic or preference-data generation.
  • Experience training LLMs on mid-to-large GPU hardware and debugging distributed training at scale.
  • Strong PyTorch fundamentals and familiarity with Hugging Face TRL, Accelerate, DeepSpeed or FSDP, and vLLM.
  • Understanding of tokenization, attention, chat templates, and common failure modes in alignment and agent training.
  • Strong communication skills and ability to work across research, product, and business teams.
  • Preferred: experience designing reward models or rule-based verifiers for reinforcement learning.
  • Preferred: experience with tool-use or agentic model training, including function calling and multi-step planning.
  • Preferred: publications or open-source contributions in LLM post-training or reinforcement learning.

Benefits

  • Health, dental, and vision care for employees and families, with 100% employee coverage
  • 401(k) plan with company matching
  • Paid time off and paid holidays
  • FSA, HSA, and commuter benefits programs
  • Team activity budget
  • Full-time position with a US work arrangement; geographic location may affect pay

Tech Stack

Categories

NewsBreak

About NewsBreak

201-500 employees

NewsBreak builds a local-news aggregation and discovery platform for U.S. consumers on web and mobile, powered by recommendation systems and NLP. It partners with publishers and independent creators and monetizes primarily through advertising for local and national businesses. Founded in 2015 and headquartered in Mountain View, California, the platform serves over 40 million monthly active users and aggregates content from more than 10,000 sources.

Contact me