1 year ago
Responsibilities
- Design and build core systems that make large training runs fast and reliable
- Build scalable distributed training infrastructure for GPU clusters
- Implement and tune parallelism and sharding strategies for evolving architectures
- Optimize distributed efficiency through topology-aware collectives, communication and computation overlap, and straggler mitigation
- Build data loading systems that eliminate I/O bottlenecks for multimodal datasets
- Develop checkpointing mechanisms that balance memory constraints with recovery needs
- Create monitoring, profiling, and debugging tools for training stability and performance
Requirements
- Hands-on experience building distributed training infrastructure with PyTorch Distributed DDP/FSDP, DeepSpeed ZeRO, or Megatron-LM TP/PP
- Experience diagnosing performance bottlenecks and failure modes, including profiling, NCCL and collective-operation issues, hangs, out-of-memory errors, and stragglers
- Understanding of hardware accelerators and networking topologies
- Experience optimizing data pipelines for machine learning workloads
- MoE training experience is preferred
- Experience with large-scale distributed training across 100 or more GPUs is preferred
- Open-source contributions to training infrastructure projects are preferred
Benefits
- Competitive base salary with equity in a unicorn-stage company
- 100% of medical, dental, and vision premiums covered for employees and dependents
- 401(k) matching up to 4% of base pay
- Unlimited PTO and company-wide Refill Days
- San Francisco and Boston are preferred, but other locations are accepted
Tech Stack
Categories
About Liquid AI
We build efficient general-purpose AI at every scale.
