4 hours ago
London, United KingdomSenior
Responsibilities
- Own optimization projects for model families, customer endpoints, and serving backends.
- Compare inference engines and recommend serving configurations for specific workloads.
- Debug model quality and performance regressions during production rollouts.
- Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, GPU utilization, quality, and cost per token.
- Deploy, configure, benchmark, and extend modern inference engines.
- Build and productionize quantization, quantization-aware training, distillation, low-bit serving, and accuracy-recovery workflows.
- Implement or integrate speculative decoding, draft-model approaches, KV-cache optimization, prefix caching, chunked prefill, continuous batching, and disaggregated prefill/decode serving.
- Build reproducible benchmark harnesses covering latency, throughput, GPU memory, reliability, and cost metrics.
- Partner with GPU kernel and platform engineers to diagnose bottlenecks across model, runtime, scheduler, gateway, and cluster layers.
- Write design documents, performance reports, rollout plans, and customer-facing technical explanations.
Requirements
- Strong Python and PyTorch engineering skills.
- Hands-on experience deploying or optimizing LLM, VLM, or high-throughput transformer inference systems.
- Practical knowledge of at least one modern inference stack, such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, or KServe.
- Strong understanding of transformer inference bottlenecks, including KV cache, attention, memory bandwidth, batching, parallelism, and long-context serving.
- Ability to reason quantitatively about latency, throughput, quality, utilization, and cost tradeoffs.
- Strong communication and cross-functional collaboration skills.
- Preferred experience with quantization-aware training, post-training quantization, FP8, INT8, INT4, NVFP4, MXFP4, AWQ, GPTQ, SmoothQuant, distillation, speculative decoding, EAGLE, Medusa, multi-token prediction, or other inference acceleration methods.
- Preferred experience with agentic workloads, tool calling, structured outputs, streaming APIs, high concurrency, and multi-step orchestration.
- CUDA or Triton familiarity is preferred.
- Open-source contributions to vLLM, SGLang, TensorRT-LLM, FlashInfer, LMCache, PyTorch, Triton, Ray, KServe, or related projects are preferred.
Benefits
- Competitive compensation.
- Career growth and learning opportunities.
- Flexibility and ownership.
- Collaborative and innovative culture.
- Opportunity to work on impactful AI projects.
- International environment with talented teams.
Categories
About Nebius
Nebius builds a full-stack AI cloud offering GPU compute, storage, and tools for training and deploying ML models for startups, enterprises, and research labs. It sells consumption-based cloud infrastructure (IaaS/PaaS) and managed services tailored to generative AI workloads, including large-scale model training and inference. The company is headquartered in Amsterdam and operates as an independent provider.
