Lightning AI

Senior Research Engineer, LLM Training & Post-Training

Lightning AI
Apply
20 hours ago
Remote, Worldwide +3 moreSenior

Base Salary

$165k - $310k/yr

Responsibilities

  • Design, build, and optimize large language model training and post-training pipelines.
  • Improve model quality through continued pretraining, supervised fine-tuning, preference optimization, reinforcement learning, evaluation, and experimentation.
  • Build and improve PyTorch-based training infrastructure, tooling, and developer workflows.
  • Optimize distributed training across multi-GPU environments for throughput, memory efficiency, scalability, and GPU utilization.
  • Investigate convergence, instability, communication overhead, and performance bottlenecks in model training.
  • Design evaluation methodologies, benchmark models, analyze failure modes, and guide improvements through experimentation.
  • Work with customers to understand real-world workloads and translate those learnings into research platform improvements.
  • Partner with research, infrastructure, and platform engineering teams to build production-ready AI systems.
  • Contribute to open-source projects through features, tooling improvements, documentation, and community engagement.

Requirements

  • Significant experience training, fine-tuning, evaluating, and optimizing transformer-based language models with PyTorch.
  • Experience with continued pretraining, supervised fine-tuning, reinforcement learning from human feedback, preference optimization including DPO, PPO, or GRPO, reward modeling, or similar techniques.
  • Strong understanding of distributed training and multi-GPU systems, including improving performance, scalability, or efficiency.
  • Strong software engineering fundamentals and experience building production-quality Python software and research tooling.
  • Experience designing experiments, evaluating model performance, and debugging complex training or optimization issues.
  • Ability to collaborate across research, product, infrastructure, and customer-facing engagements.
  • Master's degree, PhD, or equivalent industry experience in machine learning, artificial intelligence, computer science, or a related field.
  • Experience with DeepSpeed, FSDP, Megatron-LM, Hugging Face Transformers, TRL, PEFT, Lightning Fabric, or similar training frameworks is preferred.
  • Experience with CUDA, Triton, vLLM, SGLang, TensorRT, or similar AI systems and performance optimization technologies is preferred.
  • GPU performance optimization, mixed precision, memory optimization, distributed training optimization, open-source contributions, research publications, production AI platforms, startup experience, or highly cross-functional engineering experience is preferred.

Benefits

  • Comprehensive medical, dental, and vision coverage for employees and eligible dependents.
  • Meaningful equity through RSUs, retirement savings contributions, unlimited PTO, company holidays, and floating holidays.
  • Two-week company-wide winter break, paid parental and family leave, professional development allowance, wellness and work-from-home stipends, and a paid four-week sabbatical after four years.
  • Hybrid work with a minimum of two in-office days per week in San Francisco, Seattle, New York City, or London; fully remote work may be considered outside office hubs.
  • Flexible schedules, occasional team and company offsites, and complimentary meals at office hubs.

Tech Stack

Hugging Face TransformersPythonPyTorch

Categories

AI ResearchML Engineering
Lightning AI

About Lightning AI

51-200 employees

The AI development platform - From idea to AI, Lightning fast ⚡️. Code together. Prototype. Train on GPUs. Scale. Serve. From your browser - with zero setup. AI Studio is your laptop on the cloud. Zero setup. Always ready. Persistent storage and environments. Code on CPU. Debug on GPU. Scale to multi-node. Run sweeps, jobs and more. Scale models with PyTorch Lightning, Fabric, Lit-GPT, torchmetrics and more.