Member of Technical Staff, Reinforcement Learning
The Inception Company6 months ago
San Mateo, CA, USAMid Level
Responsibilities
- Design, develop, and optimize RL training pipelines for diffusion-based LLMs using PPO, DPO, RLHF, and novel approaches.
- Build and iterate on reward models, reward-shaping strategies, and reward-quality evaluation methods.
- Implement fine-tuning and scaling techniques for generative AI models.
- Develop data preprocessing pipelines, model evaluation workflows, and alignment methods for enterprise use cases.
- Research techniques for controlled text generation and constraint satisfaction.
- Improve the stability, efficiency, and reproducibility of distributed RL workloads.
Requirements
- Bachelor's, master's, or PhD in Computer Science or a related field, or equivalent experience.
- At least 2 years of experience working on ML projects using PyTorch or an equivalent framework.
- Strong familiarity with transformers and core LLM concepts, including autoregressive pretraining, instruction tuning, and in-context learning.
- Hands-on experience with RLHF, PPO, DPO, or related post-training methods.
- Familiarity with training and inference in diffusion models.
- Experience training deep learning models at scale in distributed computing environments.
- Preferred qualifications include extensive experience training transformer-based language models from scratch, designing reward models or preference-learning systems, and knowledge of mixed precision, gradient accumulation, optimization theory, and neural network architecture design.
- Experience with LLM serving frameworks such as vLLM, SGLang, or TensorRT is preferred.
Tech Stack
Categories
AI ResearchML Engineering