
Research Engineer
Lightning AI4 months ago
London, United Kingdom +3 moreSenior
Base Salary
$180k - $250k/yr
Responsibilities
- Develop graph-level, kernel-level, and system-level optimizations for deep learning training and inference workloads.
- Advance the Thunder compiler through optimization passes, graph transformations, and integration hooks.
- Build clean APIs, automated tooling, and integrations with PyTorch Lightning so optimizations are accessible to end users.
- Design and implement profiling and debugging tools to identify execution bottlenecks and guide optimization strategies.
- Collaborate with hardware vendors and ecosystem partners to support NVIDIA, AMD, TPU, and specialized accelerator backends.
- Contribute features, documentation, and community support to open-source projects.
- Engage with researchers, engineers, and external contributors on performance tuning and Thunder adoption.
- Coordinate with product and engineering teams to align compiler and optimization work with product goals.
Requirements
- Bachelor’s degree in Computer Science, Engineering, or a related field.
- Strong expertise with deep learning frameworks such as PyTorch.
- Hands-on experience with model optimization techniques including graph-level optimizations, quantization, pruning, mixed precision, or memory-efficient training.
- Knowledge of distributed systems and parallelism strategies including data, model, and pipeline parallelism, checkpointing, and elastic scaling.
- Familiarity with API design, robust tooling, testing, and CI/CD for performance-sensitive systems.
- Strong collaboration and communication skills for working with research, engineering, and external contributors.
- Experience with CUDA, Triton, or other GPU programming models is preferred.
- Deep learning compiler internals or performance-critical software experience is preferred.
- Open-source contributions in ML, HPC, or compiler domains are preferred.
- A Master’s or PhD in machine learning, compilers, or systems is highly preferred.
Benefits
- Minimum two in-office days per week in a hub office in New York City, San Francisco, Seattle, or London, with occasional team and company offsites.
- Comprehensive medical, dental, and vision coverage in the U.S.; private medical and dental insurance in the U.K.
- Retirement and financial wellness support in the U.S.; pension contribution in the U.K.
- Generous paid time off and holidays.
- Paid parental leave.
- Professional development support.
- Wellness and work-from-home stipends.
- Flexible work environment.
Tech Stack
Categories
About Lightning AI
The AI development platform - From idea to AI, Lightning fast ⚡️. Code together. Prototype. Train on GPUs. Scale. Serve. From your browser - with zero setup. AI Studio is your laptop on the cloud. Zero setup. Always ready. Persistent storage and environments. Code on CPU. Debug on GPU. Scale to multi-node. Run sweeps, jobs and more. Scale models with PyTorch Lightning, Fabric, Lit-GPT, torchmetrics and more.