
Member of Technical Staff - GPU Infrastructure
Prime Intellect2 months ago
Remote, United States or San Francisco, CA, USAMid Level
Base Salary
$150k - $300k/yr
Responsibilities
- Partner with customers to understand workload requirements and design GPU cluster architectures.
- Create technical proposals and capacity plans for clusters ranging from 100 to 10,000+ GPUs.
- Develop deployment strategies for LLM training, inference, and HPC workloads.
- Present architectural recommendations to technical and executive stakeholders.
- Deploy and configure SLURM and Kubernetes for distributed GPU workloads.
- Implement and optimize InfiniBand, RoCE, and NVLink networking.
- Optimize GPU utilization, memory management, inter-node communication, and system performance.
- Configure Lustre, BeeGFS, and GPFS parallel filesystems for high-performance I/O.
- Tune Linux kernel parameters and CUDA configurations.
- Act as the primary technical escalation point for customer infrastructure issues.
- Diagnose and resolve hardware, driver, networking, and software problems across the full stack.
- Implement monitoring, alerting, and automated remediation systems.
- Provide 24/7 on-call support for critical customer deployments.
- Create runbooks and documentation for customer operations teams.
Requirements
- At least 3 years of hands-on experience with GPU clusters and HPC environments.
- Deep production expertise with SLURM and Kubernetes in GPU settings.
- Proven experience configuring and troubleshooting InfiniBand.
- Strong understanding of NVIDIA GPU architecture, the CUDA ecosystem, and GPU drivers.
- Experience with Ansible and Terraform for infrastructure automation.
- Proficiency in Python, Bash, and systems programming.
- A track record of customer-facing technical leadership.
- Experience with NVIDIA driver installation and troubleshooting, CUDA, Fabric Manager, and DCGM.
- Experience configuring Docker, Containerd, and Enroot for GPU workloads.
- Knowledge of Linux kernel tuning, performance optimization, and AI workload network topology design.
- Understanding of power and cooling requirements for high-density GPU deployments.
- Experience with 1,000+ GPU deployments is preferred.
- NVIDIA DGX, HGX, or SuperPOD certification is preferred.
- Experience with PyTorch FSDP, DeepSpeed, Megatron-LM, ML framework optimization, or profiling is preferred.
- Experience with AMD MI300 or Intel Gaudi accelerators is preferred.
- Contributions to open-source HPC or AI infrastructure projects are preferred.
Benefits
- Cash compensation of $150,000-$300,000 plus equity incentives.
- 24/7 on-call support responsibility for critical customer deployments.
- Direct collaboration with customers and the engineering team on large-scale AI infrastructure.