Nebius

Senior ML Engineer (Token Factory)

Nebius
Apply
3 days ago
Remote, EMEA +4 moreSenior

Responsibilities

  • Identify and resolve LLM inference bottlenecks to improve production speed and cost-per-token at scale.
  • Optimize inference for dense, mixture-of-experts, autoregressive, and parallel LLM architectures.
  • Develop novel speculative decoding approaches and contribute to open-source inference engines.
  • Design and productionize low-precision FP8 and NVFP4/MXFP4 training and inference pipelines.
  • Profile GPU workloads and optimize GPU memory, compute, throughput, and latency.
  • Apply sharding strategies, custom kernels, and hardware features to large neural network training.
  • Build and maintain production-quality software using engineering practices such as version control, unit testing, and CI/CD.

Requirements

  • Profound understanding of machine learning theory and transformer architecture.
  • Experience profiling GPU workloads with Nsight, PyTorch profiler, or similar tools.
  • Understanding of GPU memory hierarchy and compute/memory tradeoffs.
  • Familiarity with MHA, RoPE, KV-cache, Flash Attention, and quantisation.
  • Understanding of performance considerations in large neural network training, including sharding strategies, custom kernels, and hardware features.
  • Strong software engineering skills, primarily in Python.
  • Deep experience with modern deep learning frameworks.
  • Proficiency with CI/CD, version control, and unit testing.
  • Strong communication and leadership abilities.
  • Preferred experience contributing to open-source inference engines such as vLLM, SGLang, or TensorRT-LLM.
  • Preferred experience with Triton, Cute, CUTLASS, or CUDA.
  • Preferred experience developing large distributed systems or high-load web services.
  • Open-source projects and experience delivering products in a dynamic startup-like environment are preferred.
  • Excellent English writing, articulation, and communication skills.

Benefits

  • Competitive compensation.
  • Career growth and learning opportunities.
  • Flexibility, ownership, and an opportunity to shape the future of AI.
  • Collaborative, innovative, and international working environment.
  • Opportunity to work on impactful AI projects with talented teams.

Tech Stack

Categories

Nebius

About Nebius

1,001-5,000 employees

The Nebius AI Cloud brings powerful full-stack infrastructure for AI developers and practitioners across startups, enterprises and science institutes to build and deploy generative AI applications and rapidly deliver scientific breakthroughs by training and running ML models within a secure, high-performance, and cost-optimized cloud environment.

Contact me