Nebius

ML Infrastructure Engineer

Nebius
Apply
4 months ago
Remote, United States or Remote, EMEASenior

Responsibilities

  • Profile and analyze GPU performance at the system and kernel levels with hardware and development teams.
  • Evaluate and compare GPU performance across platforms, architectures, and software stacks including CUDA and ROCm.
  • Debug and optimize machine-learning workloads on GPU hardware and resolve performance bottlenecks.
  • Conduct acceptance testing for new GPU clusters to verify performance, stability, and compatibility for AI workloads.
  • Experiment with GPU system configurations, interconnect strategies, and system-level optimizations to assess performance and scalability.
  • Develop tools and dashboards to visualize performance metrics, bottlenecks, and trends.
  • Contribute to internal tooling, frameworks, and engineering best practices.

Requirements

  • Profound understanding of machine-learning theoretical foundations.
  • Deep understanding of performance considerations for large neural-network training and inference, including parallelism, offloading, custom kernels, hardware features, attention optimizations, and dynamic batching.
  • Deep experience with PyTorch, JAX, Megatron-LM, and TensorRT-LLM.
  • Good understanding of CUDA, NCCL, drivers, and relevant GPU libraries.
  • Familiarity with Docker and Kubernetes.
  • Strong communication skills and ability to work independently.
  • Familiarity with vLLM, SGLang, and TensorRT is preferred.
  • Experience with Python and performance-profiling tools such as Nsight, nvprof, and perf is preferred.
  • Familiarity with AWS, GCP, and Azure ML is preferred.
  • Contributions to open-source machine-learning benchmarking tools are preferred.
  • Applicants must be authorized to work in the country in which they apply and provide proof of employment eligibility.

Benefits

  • Competitive compensation (amount not stated)
  • Career growth and learning opportunities
  • Flexible work-life balance
  • Collaborative and innovative culture
  • Opportunity to work on impactful AI projects
  • International environment and talented teams
Nebius

About Nebius

1,001-5,000 employees

Nebius builds a full-stack AI cloud offering GPU compute, storage, and tools for training and deploying ML models for startups, enterprises, and research labs. It sells consumption-based cloud infrastructure (IaaS/PaaS) and managed services tailored to generative AI workloads, including large-scale model training and inference. The company is headquartered in Amsterdam and operates as an independent provider.

Contact me