1 month ago
Base Salary
$150k - $220k/yr
Responsibilities
- Own versioned OS images, kernel tuning, GPU and NIC driver stacks, and automated pipelines for bringing nodes into production.
- Build automated node acceptance and burn-in suites using GPU diagnostics, communication tests, bandwidth and topology checks, and HPL.
- Execute rolling kernel, driver, and toolkit upgrades while maintaining fleet consistency and compatibility across CUDA, ROCm, and ML frameworks.
- Automate unhealthy-node detection and cordon, drain, reboot, re-image, and hardware-escalation workflows.
- Manage infrastructure configuration through Ansible or SaltStack, Git-based reviews, CI validation, and canary rollouts.
- Build reproducible provisioning and image pipelines using tools such as PXE, MaaS, Packer, or similar technologies.
- Operate GPU-enabled Kubernetes for inference and Slurm or Run:AI for multi-node training workloads.
- Maintain container images, registries, NVIDIA and ROCm container stacks, and optimized PyTorch and JAX environments.
- Deploy and debug NVIDIA and AMD driver/runtime stacks, kernel modules, GPUDirect, RDMA software, and distributed GPU communication.
- Tune distributed workloads across NVLink, NVSwitch, InfiniBand, and RoCE fabrics and monitor GPU utilization and cluster efficiency.
- Develop Python and Bash tooling for cluster operations, health reporting, workflow automation, and observability.
- Troubleshoot NCCL hangs, CUDA memory leaks, ROCm crashes, straggler nodes, and unexplained throughput degradation.
Requirements
- 5+ years of systems or infrastructure engineering experience with significant GPU cluster, HPC, or large-scale ML infrastructure experience.
- Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field.
- Deep Linux internals expertise, including kernel modules, DKMS, systemd, cgroups, NUMA, and performance tuning.
- Hands-on experience with NVIDIA CUDA and/or AMD ROCm driver and runtime stacks, including kernel-level debugging.
- Production Kubernetes experience with GPU workloads and working knowledge of Slurm or Run:AI, or equivalent depth in the reverse order.
- Strong Ansible or SaltStack configuration-management experience with Git-based, code-reviewed infrastructure workflows.
- Experience with Packer, MaaS, Foreman, Terraform, or similar automated provisioning and image tools.
- Client-side experience with Lustre, GPFS, or Weka distributed filesystems and checkpoint I/O optimization.
- Container experience with Docker or containerd and the NVIDIA Container Toolkit or ROCm equivalent.
- Proficiency in Python and Bash.
- Working knowledge of NCCL, RDMA networking, InfiniBand or RoCE, GPUDirect, and PyTorch or JAX runtime behavior.
- Preferred experience supporting foundation-model training teams, tuning inference stacks such as vLLM, Triton Inference Server, or TensorRT-LLM, and using Nsight Systems, Nsight Compute, rocprof, perf, or eBPF.
Benefits
- Medical, dental, and vision insurance.
- 401k plan.
- Daily lunch, snacks, and beverages.
- Flexible time off.
- Competitive salary and equity.
