Nebius

Senior HPC Cluster Engineer

Nebius
Apply
4 days ago
Remote, EMEA or Madrid, SpainSenior

Responsibilities

  • Tune GPU clusters and InfiniBand networks for high-performance HPC and GPU workloads.
  • Analyze and troubleshoot GPU and InfiniBand issues and propose corrective actions.
  • Integrate new hardware and support new GPUs through Kubernetes, QEMU, and KVM software stacks.
  • Enhance automation for monitoring, fault detection, and issue resolution.
  • Configure and manage GPU devices and InfiniBand fabrics.
  • Analyze and optimize HPC workloads and system performance.

Requirements

  • 5+ years of professional experience in system-level software development focused on performance optimization and low-level programming.
  • 3+ years of hands-on Linux administration, troubleshooting, and performance tuning experience.
  • In-depth understanding of server architecture, PCIe devices, NICs, the Linux OS/kernel, and HPC systems.
  • Strong proficiency in one or more of C, C++, Go, or Python.
  • Preferred experience with GPU end-to-end testing in cluster environments using InfiniBand.
  • Preferred experience optimizing HPC workloads such as simulations, data analysis, and AI/ML workloads.
  • Familiarity with RDMA, RoCE, InfiniBand, software-defined networking, and HPC cluster networking.
  • Understanding of QEMU/KVM virtualization and virtualized environments.
  • Experience with PyTorch, TensorFlow, MPI, and NCCL is a plus.
  • Applicants must be authorized to work in the country where they apply.

Benefits

  • Competitive compensation.
  • Career growth and learning opportunities.
  • Flexibility and ownership.
  • Collaborative and innovative culture.
  • Opportunity to work on impactful AI projects.
  • International environment and talented teams.
  • Applicants must provide proof of employment eligibility as a condition of hire.

Categories

Nebius

About Nebius

1,001-5,000 employees

Nebius builds a full-stack AI cloud offering GPU compute, storage, and tools for training and deploying ML models for startups, enterprises, and research labs. It sells consumption-based cloud infrastructure (IaaS/PaaS) and managed services tailored to generative AI workloads, including large-scale model training and inference. The company is headquartered in Amsterdam and operates as an independent provider.

Contact me