Liquid AI

Member of Technical Staff - Distributed Training Engineer

Liquid AI
Apply
1 year ago
Remote, Worldwide +2 moreSenior
H1B Sponsor

Responsibilities

  • Design and build core systems that make large training runs fast and reliable
  • Build scalable distributed training infrastructure for GPU clusters
  • Implement and tune parallelism and sharding strategies for evolving architectures
  • Optimize distributed efficiency through topology-aware collectives, communication and computation overlap, and straggler mitigation
  • Build data loading systems that eliminate I/O bottlenecks for multimodal datasets
  • Develop checkpointing mechanisms that balance memory constraints with recovery needs
  • Create monitoring, profiling, and debugging tools for training stability and performance

Requirements

  • Hands-on experience building distributed training infrastructure with PyTorch Distributed DDP/FSDP, DeepSpeed ZeRO, or Megatron-LM TP/PP
  • Experience diagnosing performance bottlenecks and failure modes, including profiling, NCCL and collective-operation issues, hangs, out-of-memory errors, and stragglers
  • Understanding of hardware accelerators and networking topologies
  • Experience optimizing data pipelines for machine learning workloads
  • MoE training experience is preferred
  • Experience with large-scale distributed training across 100 or more GPUs is preferred
  • Open-source contributions to training infrastructure projects are preferred

Benefits

  • Competitive base salary with equity in a unicorn-stage company
  • 100% of medical, dental, and vision premiums covered for employees and dependents
  • 401(k) matching up to 4% of base pay
  • Unlimited PTO and company-wide Refill Days
  • San Francisco and Boston are preferred, but other locations are accepted

Tech Stack

Liquid AI

About Liquid AI

51-200 employees

We build efficient general-purpose AI at every scale.