4 days ago
Remote, EMEA or Madrid, SpainSenior
Responsibilities
- Tune GPU clusters and InfiniBand networks for high-performance HPC and GPU workloads.
- Analyze and troubleshoot GPU and InfiniBand issues and propose corrective actions.
- Integrate new hardware and support new GPUs through Kubernetes, QEMU, and KVM software stacks.
- Enhance automation for monitoring, fault detection, and issue resolution.
- Configure and manage GPU devices and InfiniBand fabrics.
- Analyze and optimize HPC workloads and system performance.
Requirements
- 5+ years of professional experience in system-level software development focused on performance optimization and low-level programming.
- 3+ years of hands-on Linux administration, troubleshooting, and performance tuning experience.
- In-depth understanding of server architecture, PCIe devices, NICs, the Linux OS/kernel, and HPC systems.
- Strong proficiency in one or more of C, C++, Go, or Python.
- Preferred experience with GPU end-to-end testing in cluster environments using InfiniBand.
- Preferred experience optimizing HPC workloads such as simulations, data analysis, and AI/ML workloads.
- Familiarity with RDMA, RoCE, InfiniBand, software-defined networking, and HPC cluster networking.
- Understanding of QEMU/KVM virtualization and virtualized environments.
- Experience with PyTorch, TensorFlow, MPI, and NCCL is a plus.
- Applicants must be authorized to work in the country where they apply.
Benefits
- Competitive compensation.
- Career growth and learning opportunities.
- Flexibility and ownership.
- Collaborative and innovative culture.
- Opportunity to work on impactful AI projects.
- International environment and talented teams.
- Applicants must provide proof of employment eligibility as a condition of hire.
About Nebius
Nebius builds a full-stack AI cloud offering GPU compute, storage, and tools for training and deploying ML models for startups, enterprises, and research labs. It sells consumption-based cloud infrastructure (IaaS/PaaS) and managed services tailored to generative AI workloads, including large-scale model training and inference. The company is headquartered in Amsterdam and operates as an independent provider.
