Sciforium

GPU Cluster Engineer, Systems & Platform

Sciforium
Apply
1 month ago

Base Salary

$150k - $220k/yr

Responsibilities

  • Own versioned OS images, kernel tuning, GPU and NIC driver stacks, and automated pipelines for bringing nodes into production.
  • Build automated node acceptance and burn-in suites using GPU diagnostics, communication tests, bandwidth and topology checks, and HPL.
  • Execute rolling kernel, driver, and toolkit upgrades while maintaining fleet consistency and compatibility across CUDA, ROCm, and ML frameworks.
  • Automate unhealthy-node detection and cordon, drain, reboot, re-image, and hardware-escalation workflows.
  • Manage infrastructure configuration through Ansible or SaltStack, Git-based reviews, CI validation, and canary rollouts.
  • Build reproducible provisioning and image pipelines using tools such as PXE, MaaS, Packer, or similar technologies.
  • Operate GPU-enabled Kubernetes for inference and Slurm or Run:AI for multi-node training workloads.
  • Maintain container images, registries, NVIDIA and ROCm container stacks, and optimized PyTorch and JAX environments.
  • Deploy and debug NVIDIA and AMD driver/runtime stacks, kernel modules, GPUDirect, RDMA software, and distributed GPU communication.
  • Tune distributed workloads across NVLink, NVSwitch, InfiniBand, and RoCE fabrics and monitor GPU utilization and cluster efficiency.
  • Develop Python and Bash tooling for cluster operations, health reporting, workflow automation, and observability.
  • Troubleshoot NCCL hangs, CUDA memory leaks, ROCm crashes, straggler nodes, and unexplained throughput degradation.

Requirements

  • 5+ years of systems or infrastructure engineering experience with significant GPU cluster, HPC, or large-scale ML infrastructure experience.
  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field.
  • Deep Linux internals expertise, including kernel modules, DKMS, systemd, cgroups, NUMA, and performance tuning.
  • Hands-on experience with NVIDIA CUDA and/or AMD ROCm driver and runtime stacks, including kernel-level debugging.
  • Production Kubernetes experience with GPU workloads and working knowledge of Slurm or Run:AI, or equivalent depth in the reverse order.
  • Strong Ansible or SaltStack configuration-management experience with Git-based, code-reviewed infrastructure workflows.
  • Experience with Packer, MaaS, Foreman, Terraform, or similar automated provisioning and image tools.
  • Client-side experience with Lustre, GPFS, or Weka distributed filesystems and checkpoint I/O optimization.
  • Container experience with Docker or containerd and the NVIDIA Container Toolkit or ROCm equivalent.
  • Proficiency in Python and Bash.
  • Working knowledge of NCCL, RDMA networking, InfiniBand or RoCE, GPUDirect, and PyTorch or JAX runtime behavior.
  • Preferred experience supporting foundation-model training teams, tuning inference stacks such as vLLM, Triton Inference Server, or TensorRT-LLM, and using Nsight Systems, Nsight Compute, rocprof, perf, or eBPF.

Benefits

  • Medical, dental, and vision insurance.
  • 401k plan.
  • Daily lunch, snacks, and beverages.
  • Flexible time off.
  • Competitive salary and equity.

Tech Stack

Categories

Sciforium

About Sciforium

11-50 employees
Contact me