7 hours ago
Remote, EMEASenior
Responsibilities
- Own optimization work for specific model families, customer endpoints, or serving backends.
- Compare inference engines and recommend serving configurations for specific workloads.
- Debug model-quality and performance regressions during production rollouts.
- Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, GPU utilization, quality, and cost per token.
- Deploy, configure, benchmark, and extend inference engines such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, and NVIDIA Dynamo.
- Build and productionize model-compression workflows including quantization, quantization-aware training, distillation, low-bit serving, and accuracy recovery.
- Implement or integrate speculative decoding, draft-model approaches, KV-cache optimization, prefix caching, chunked prefill, continuous batching, and disaggregated prefill/decode serving.
- Build reproducible benchmark harnesses measuring TTFT, TPOT, tokens per second per GPU, p95/p99 latency, GPU memory, reliability, and cost per token.
- Partner with GPU kernel and platform engineers to diagnose bottlenecks across model code, kernels, runtime, scheduler, gateway, and cluster layers.
- Write design documents, performance reports, rollout plans, and customer-facing technical explanations.
Requirements
- Strong Python and PyTorch engineering skills.
- Hands-on experience deploying or optimizing LLM, VLM, or high-throughput transformer inference systems.
- Practical knowledge of at least one modern inference stack such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, or equivalent internal systems.
- Strong understanding of transformer inference bottlenecks including KV cache, attention, memory bandwidth, batching, parallelism, and long-context serving.
- Ability to reason quantitatively about latency, throughput, quality, utilization, and cost tradeoffs.
- Strong communication and collaboration skills across research, kernel, infrastructure, product, and customer teams.
- Preferred experience with quantization-aware training, post-training quantization, FP8, INT8, INT4, NVFP4, MXFP4, AWQ, GPTQ, SmoothQuant, or related techniques.
- Preferred experience with distillation, speculative decoding, EAGLE, Medusa, multi-token prediction, or other inference-acceleration methods.
- Preferred experience with agentic workloads involving tool calling, structured outputs, streaming APIs, high concurrency, and multi-step orchestration.
- CUDA or Triton familiarity is preferred.
- Open-source contributions to vLLM, SGLang, TensorRT-LLM, FlashInfer, LMCache, PyTorch, Triton, Ray, KServe, or related projects are preferred.
Benefits
- Competitive compensation.
- Career growth and learning opportunities.
- Flexibility and ownership.
- Collaborative and innovative culture.
- Opportunity to work on impactful AI projects.
- International environment with talented teams.
Categories
About Nebius
Nebius builds a full-stack AI cloud offering GPU compute, storage, and tools for training and deploying ML models for startups, enterprises, and research labs. It sells consumption-based cloud infrastructure (IaaS/PaaS) and managed services tailored to generative AI workloads, including large-scale model training and inference. The company is headquartered in Amsterdam and operates as an independent provider.
