Prime Intellect

Member of Technical Staff - GPU Infrastructure

Prime Intellect
Apply
2 months ago
Remote, United States or San Francisco, CA, USAMid Level

Base Salary

$150k - $300k/yr

Responsibilities

  • Partner with customers to understand workload requirements and design GPU cluster architectures.
  • Create technical proposals and capacity plans for clusters ranging from 100 to 10,000+ GPUs.
  • Develop deployment strategies for LLM training, inference, and HPC workloads.
  • Present architectural recommendations to technical and executive stakeholders.
  • Deploy and configure SLURM and Kubernetes for distributed GPU workloads.
  • Implement and optimize InfiniBand, RoCE, and NVLink networking.
  • Optimize GPU utilization, memory management, inter-node communication, and system performance.
  • Configure Lustre, BeeGFS, and GPFS parallel filesystems for high-performance I/O.
  • Tune Linux kernel parameters and CUDA configurations.
  • Act as the primary technical escalation point for customer infrastructure issues.
  • Diagnose and resolve hardware, driver, networking, and software problems across the full stack.
  • Implement monitoring, alerting, and automated remediation systems.
  • Provide 24/7 on-call support for critical customer deployments.
  • Create runbooks and documentation for customer operations teams.

Requirements

  • At least 3 years of hands-on experience with GPU clusters and HPC environments.
  • Deep production expertise with SLURM and Kubernetes in GPU settings.
  • Proven experience configuring and troubleshooting InfiniBand.
  • Strong understanding of NVIDIA GPU architecture, the CUDA ecosystem, and GPU drivers.
  • Experience with Ansible and Terraform for infrastructure automation.
  • Proficiency in Python, Bash, and systems programming.
  • A track record of customer-facing technical leadership.
  • Experience with NVIDIA driver installation and troubleshooting, CUDA, Fabric Manager, and DCGM.
  • Experience configuring Docker, Containerd, and Enroot for GPU workloads.
  • Knowledge of Linux kernel tuning, performance optimization, and AI workload network topology design.
  • Understanding of power and cooling requirements for high-density GPU deployments.
  • Experience with 1,000+ GPU deployments is preferred.
  • NVIDIA DGX, HGX, or SuperPOD certification is preferred.
  • Experience with PyTorch FSDP, DeepSpeed, Megatron-LM, ML framework optimization, or profiling is preferred.
  • Experience with AMD MI300 or Intel Gaudi accelerators is preferred.
  • Contributions to open-source HPC or AI infrastructure projects are preferred.

Benefits

  • Cash compensation of $150,000-$300,000 plus equity incentives.
  • 24/7 on-call support responsibility for critical customer deployments.
  • Direct collaboration with customers and the engineering team on large-scale AI infrastructure.

Categories

Solutions Engineering
Prime Intellect

About Prime Intellect

51-200 employees
Contact me