Nebius

Senior Machine Learning Engineer, LLM Inference Optimization

Nebius
Apply
2 months ago
Palo Alto, CA, USASenior

Base Salary

$195k - $262k/yr

Responsibilities

  • Own optimization work for specified model families, customer endpoints, or serving backends.
  • Compare inference engines and recommend serving configurations for specific workloads.
  • Debug model quality and performance regressions during production rollouts.
  • Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, GPU utilization, quality, and cost per token.
  • Deploy, configure, benchmark, and extend inference engines including vLLM, SGLang, TensorRT-LLM, Triton Inference Server, and NVIDIA Dynamo.
  • Build and productionize compression workflows covering quantization, quantization-aware training, distillation, low-bit serving, and accuracy recovery.
  • Implement or integrate speculative decoding, draft-model approaches, KV-cache optimization, prefix caching, chunked prefill, continuous batching, and disaggregated prefill/decode serving.
  • Build reproducible benchmarks for TTFT, TPOT, tokens per second per GPU, p95/p99 latency, GPU memory, reliability, and cost per token.
  • Partner with GPU kernel and platform engineers to diagnose bottlenecks across model code, kernels, runtime, scheduler, gateway, and cluster layers.
  • Write design documents, performance reports, rollout plans, and customer-facing technical explanations.

Requirements

  • Strong Python and PyTorch engineering skills.
  • Hands-on experience deploying or optimizing LLM, VLM, or high-throughput transformer inference systems.
  • Practical knowledge of at least one modern inference stack, such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, or equivalent internal systems.
  • Strong understanding of transformer inference bottlenecks, including KV cache, attention, memory bandwidth, batching, parallelism, and long-context serving.
  • Ability to reason quantitatively about latency, throughput, quality, utilization, and cost tradeoffs.
  • Strong communication and collaboration skills across research, kernel, infrastructure, product, and customer teams.
  • Experience with quantization-aware training, post-training quantization, FP8, INT8, INT4, NVFP4, MXFP4, AWQ, GPTQ, SmoothQuant, or related techniques is preferred.
  • Experience with distillation, speculative decoding, EAGLE, Medusa, multi-token prediction, agentic workloads, CUDA, Triton, or open-source contributions to relevant inference projects is preferred.

Benefits

  • 100% company-paid medical, dental, and vision coverage for employees and families.
  • 401(k) plan with up to 4% company match and immediate vesting.
  • 20 weeks of paid parental leave for primary caregivers and 12 weeks for secondary caregivers.
  • Up to $85 per month in remote work reimbursement for mobile and internet.
  • Company-paid short-term disability, long-term disability, and life insurance.
  • Career growth and learning opportunities, flexibility and ownership, and an international collaborative environment.
Nebius

About Nebius

1,001-5,000 employees

The Nebius AI Cloud brings powerful full-stack infrastructure for AI developers and practitioners across startups, enterprises and science institutes to build and deploy generative AI applications and rapidly deliver scientific breakthroughs by training and running ML models within a secure, high-performance, and cost-optimized cloud environment.

Contact me