1 month ago
Remote, United StatesStaff+
Responsibilities
- Design and operate distributed training systems for large neural networks across GPU clusters.
- Optimize multi-node and multi-GPU execution, GPU utilization, throughput, memory usage, and networking performance.
- Build and optimize GPU cluster orchestration with Slurm, Kubernetes, Ray, and RunAI.
- Debug distributed communication using NCCL, RDMA, InfiniBand, and NVLink.
- Scale training with PyTorch Distributed, Megatron-LM, and DeepSpeed, including launch configurations, failure recovery, and performance tuning.
- Apply activation checkpointing, ZeRO Stages 1–3, and offload strategies to increase model and batch scale.
- Improve training stability and fault tolerance and partner with research and applied ML teams to productionize large-model training pipelines.
Requirements
- Deep hands-on experience with distributed systems or ML systems.
- Production experience running large-scale workloads on GPU clusters.
- Production experience with PyTorch distributed training.
- Strong understanding of data, tensor, and pipeline parallelism strategies.
- Low-level understanding of GPU communication and networking.
- Experience with GPU orchestration tools including Slurm, Kubernetes, Ray, and RunAI.
- Experience with NCCL, RDMA, InfiniBand, NVLink, PyTorch Distributed, Megatron-LM, DeepSpeed, activation checkpointing, and ZeRO offload techniques.
- Experience as an ML Systems Engineer, Distributed Systems Engineer, AI Infrastructure Engineer, or HPC Engineer transitioning into ML is relevant.
- Experience with large language models or foundation models is a strong plus.