2 hours ago
London, United KingdomSenior
Responsibilities
- Profile, benchmark, and optimize large-scale LLM inference workloads across compute infrastructure.
- Deploy models efficiently across GPU architectures and adapt inference systems as hardware evolves.
- Design and implement inference optimizations while maintaining output quality.
- Develop reference implementations, libraries, and tooling for efficient and reliable NLP workloads.
- Collaborate with research, infrastructure, engineering, systems, architecture, and platform teams.
- Influence long-term compute stack and platform decisions through performance analysis and optimized solutions.
Requirements
- Bachelor’s, Master’s, or PhD in computer science, or equivalent experience.
- Proven experience profiling, benchmarking, and optimizing large-scale LLM inference workloads.
- Scientific, evidence-led approach to performance optimization using rigorous benchmarking and reproducible measurement.
- Deep understanding of transformer inference, including prefill versus decode, KV-cache behavior, attention variants, and performance bottlenecks.
- Hands-on experience with LLM serving frameworks such as vLLM, SGLang, TensorRT-LLM, or TGI and the PyTorch ecosystem.
- Experience with quantization, speculative decoding, and model parallelism across modern GPU architectures.
- Strong software engineering skills, including Python, CUDA, and building reliable systems for machine learning workloads.
- Strong communication and cross-functional collaboration skills.
Benefits
- Highly competitive compensation plus an annual discretionary bonus.
- Lunch provided via Just Eat for Business and a dedicated barista bar.
- 35 days of annual leave.
- 9% company pension contributions.
- Informal dress code and excellent work/life balance.
- Comprehensive healthcare and life assurance.
- Cycle-to-work scheme and monthly company events.
- Based at the company's London headquarters.
