Nebius

Senior Machine Learning Engineer, LLM Inference Optimization

Nebius
Apply
4 hours ago
London, United KingdomSenior

Responsibilities

  • Own optimization projects for model families, customer endpoints, and serving backends.
  • Compare inference engines and recommend serving configurations for specific workloads.
  • Debug model quality and performance regressions during production rollouts.
  • Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, GPU utilization, quality, and cost per token.
  • Deploy, configure, benchmark, and extend modern inference engines.
  • Build and productionize quantization, quantization-aware training, distillation, low-bit serving, and accuracy-recovery workflows.
  • Implement or integrate speculative decoding, draft-model approaches, KV-cache optimization, prefix caching, chunked prefill, continuous batching, and disaggregated prefill/decode serving.
  • Build reproducible benchmark harnesses covering latency, throughput, GPU memory, reliability, and cost metrics.
  • Partner with GPU kernel and platform engineers to diagnose bottlenecks across model, runtime, scheduler, gateway, and cluster layers.
  • Write design documents, performance reports, rollout plans, and customer-facing technical explanations.

Requirements

  • Strong Python and PyTorch engineering skills.
  • Hands-on experience deploying or optimizing LLM, VLM, or high-throughput transformer inference systems.
  • Practical knowledge of at least one modern inference stack, such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, or KServe.
  • Strong understanding of transformer inference bottlenecks, including KV cache, attention, memory bandwidth, batching, parallelism, and long-context serving.
  • Ability to reason quantitatively about latency, throughput, quality, utilization, and cost tradeoffs.
  • Strong communication and cross-functional collaboration skills.
  • Preferred experience with quantization-aware training, post-training quantization, FP8, INT8, INT4, NVFP4, MXFP4, AWQ, GPTQ, SmoothQuant, distillation, speculative decoding, EAGLE, Medusa, multi-token prediction, or other inference acceleration methods.
  • Preferred experience with agentic workloads, tool calling, structured outputs, streaming APIs, high concurrency, and multi-step orchestration.
  • CUDA or Triton familiarity is preferred.
  • Open-source contributions to vLLM, SGLang, TensorRT-LLM, FlashInfer, LMCache, PyTorch, Triton, Ray, KServe, or related projects are preferred.

Benefits

  • Competitive compensation.
  • Career growth and learning opportunities.
  • Flexibility and ownership.
  • Collaborative and innovative culture.
  • Opportunity to work on impactful AI projects.
  • International environment with talented teams.

Tech Stack

Categories

Nebius

About Nebius

1,001-5,000 employees

Nebius builds a full-stack AI cloud offering GPU compute, storage, and tools for training and deploying ML models for startups, enterprises, and research labs. It sells consumption-based cloud infrastructure (IaaS/PaaS) and managed services tailored to generative AI workloads, including large-scale model training and inference. The company is headquartered in Amsterdam and operates as an independent provider.

Contact me