
Senior Research Engineer, LLM Training & Post-Training
Lightning AI20 hours ago
Remote, Worldwide +3 moreSenior
Base Salary
$165k - $310k/yr
Responsibilities
- Design, build, and optimize large language model training and post-training pipelines.
- Improve model quality through continued pretraining, supervised fine-tuning, preference optimization, reinforcement learning, evaluation, and experimentation.
- Build and improve PyTorch-based training infrastructure, tooling, and developer workflows.
- Optimize distributed training across multi-GPU environments for throughput, memory efficiency, scalability, and GPU utilization.
- Investigate convergence, instability, communication overhead, and performance bottlenecks in model training.
- Design evaluation methodologies, benchmark models, analyze failure modes, and guide improvements through experimentation.
- Work with customers to understand real-world workloads and translate those learnings into research platform improvements.
- Partner with research, infrastructure, and platform engineering teams to build production-ready AI systems.
- Contribute to open-source projects through features, tooling improvements, documentation, and community engagement.
Requirements
- Significant experience training, fine-tuning, evaluating, and optimizing transformer-based language models with PyTorch.
- Experience with continued pretraining, supervised fine-tuning, reinforcement learning from human feedback, preference optimization including DPO, PPO, or GRPO, reward modeling, or similar techniques.
- Strong understanding of distributed training and multi-GPU systems, including improving performance, scalability, or efficiency.
- Strong software engineering fundamentals and experience building production-quality Python software and research tooling.
- Experience designing experiments, evaluating model performance, and debugging complex training or optimization issues.
- Ability to collaborate across research, product, infrastructure, and customer-facing engagements.
- Master's degree, PhD, or equivalent industry experience in machine learning, artificial intelligence, computer science, or a related field.
- Experience with DeepSpeed, FSDP, Megatron-LM, Hugging Face Transformers, TRL, PEFT, Lightning Fabric, or similar training frameworks is preferred.
- Experience with CUDA, Triton, vLLM, SGLang, TensorRT, or similar AI systems and performance optimization technologies is preferred.
- GPU performance optimization, mixed precision, memory optimization, distributed training optimization, open-source contributions, research publications, production AI platforms, startup experience, or highly cross-functional engineering experience is preferred.
Benefits
- Comprehensive medical, dental, and vision coverage for employees and eligible dependents.
- Meaningful equity through RSUs, retirement savings contributions, unlimited PTO, company holidays, and floating holidays.
- Two-week company-wide winter break, paid parental and family leave, professional development allowance, wellness and work-from-home stipends, and a paid four-week sabbatical after four years.
- Hybrid work with a minimum of two in-office days per week in San Francisco, Seattle, New York City, or London; fully remote work may be considered outside office hubs.
- Flexible schedules, occasional team and company offsites, and complimentary meals at office hubs.
Categories
AI ResearchML Engineering
About Lightning AI
The AI development platform - From idea to AI, Lightning fast ⚡️. Code together. Prototype. Train on GPUs. Scale. Serve. From your browser - with zero setup. AI Studio is your laptop on the cloud. Zero setup. Always ready. Persistent storage and environments. Code on CPU. Debug on GPU. Scale to multi-node. Run sweeps, jobs and more. Scale models with PyTorch Lightning, Fabric, Lit-GPT, torchmetrics and more.