Nebius

Senior ML Engineer (Token Factory)

Nebius
Apply
3 days ago
Amsterdam, NetherlandsSenior

Responsibilities

  • Develop and optimize low-level kernels and runtime components for AI inference.
  • Improve the performance of inference engines on GPU platforms.
  • Profile and debug system-level and hardware-level performance issues.
  • Integrate support for Hopper, Blackwell, and Rubin hardware architectures.
  • Collaborate with ML and backend teams to optimize end-to-end execution.

Requirements

  • Strong proficiency in C++ or expertise in GPU programming focused on low-level, high-performance coding and memory management.
  • Experience with GPU programming or systems-level software development, such as operating-system internals, kernel modules, or device drivers.
  • Hands-on experience using profiling and debugging tools to identify CPU and GPU performance issues and optimize code based on findings.
  • Solid understanding of CPU/GPU architecture and memory hierarchy.
  • Preferred: experience with CUDA, ROCm, CUTLASS, Cute, ThunderKittens, Triton, Pallas, or Mosaic GPU.
  • Preferred: familiarity with ML inference runtimes such as TensorRT or TVM.
  • Preferred: knowledge of Linux internals, drivers, or compiler toolchains.
  • Preferred: experience with perf, VTune, Nsight, or ROCm profiler.
  • Preferred: familiarity with inference engines such as vLLM, sglang, or TGI.
  • Coding interviews are part of the hiring process.
  • Applicants must be authorized to work in the country of application and provide proof of employment eligibility.

Benefits

  • Competitive compensation.
  • Career growth and learning opportunities.
  • Flexibility and ownership.
  • Collaborative and innovative culture.
  • Opportunity to work on impactful AI projects.
  • International environment with talented teams.

Tech Stack

Categories

Nebius

About Nebius

1,001-5,000 employees

The Nebius AI Cloud brings powerful full-stack infrastructure for AI developers and practitioners across startups, enterprises and science institutes to build and deploy generative AI applications and rapidly deliver scientific breakthroughs by training and running ML models within a secure, high-performance, and cost-optimized cloud environment.

Contact me