Nebius

Senior Systems Software Engineer, GPU Compute

Nebius
Apply
3 days ago
Remote, United StatesSenior

Base Salary

$170k - $300k/yr

Responsibilities

  • Tune the performance of GPU clusters and InfiniBand networks in HPC and GPU-based environments.
  • Analyze and troubleshoot root causes involving GPUs and InfiniBand networks and propose corrective actions.
  • Integrate new hardware, including new GPUs, through Kubernetes, QEMU, and KVM software stacks.
  • Enhance automation for proactive monitoring, issue detection, and resolution in GPU and InfiniBand environments.
  • Configure and manage GPU devices and InfiniBand fabrics.
  • Participate in coding interviews as part of the hiring process.

Requirements

  • 5+ years of professional experience in system-level software development focused on performance optimization and low-level programming.
  • 3+ years of hands-on experience with Linux administration, troubleshooting, and performance tuning.
  • In-depth understanding of server architecture, PCIe devices, NICs, the Linux OS/kernel, and HPC systems.
  • Strong proficiency in one or more of C/C++, Go, or Python.
  • Experience with GPU end-to-end testing in cluster environments using InfiniBand is preferred.
  • Experience optimizing HPC workloads such as simulations, data analysis, or AI/ML workloads is preferred.
  • Familiarity with RDMA, RoCE, and InfiniBand protocols is preferred.
  • Background in Software-Defined Networking and HPC cluster networking is preferred.
  • Understanding of QEMU/KVM virtualization and virtualized environments is preferred.
  • Experience with PyTorch and TensorFlow integration with HPC systems is preferred.
  • Familiarity with MPI and NCCL for distributed computing is preferred.
  • Applicants must be authorized to work in the country where they apply.

Benefits

  • Competitive compensation ranging from $170k-$300k plus equity based on experience.
  • Career growth and learning opportunities.
  • Flexibility and ownership.
  • Collaborative and innovative culture.
  • Opportunity to work on impactful AI projects.
  • International environment with talented teams.

Categories

Nebius

About Nebius

1,001-5,000 employees

The Nebius AI Cloud brings powerful full-stack infrastructure for AI developers and practitioners across startups, enterprises and science institutes to build and deploy generative AI applications and rapidly deliver scientific breakthroughs by training and running ML models within a secure, high-performance, and cost-optimized cloud environment.

Contact me