Together AI

Research Engineer, Large-Scale Training

Together AI
Apply
2 months ago

Base Salary

$200k - $290k/yr

Responsibilities

  • Design, implement, and optimize core components of large-scale training infrastructure.
  • Integrate new model architectures, validate training correctness and convergence, and optimize production fine-tuning workloads.
  • Profile distributed training workloads and remove compute, memory, and communication bottlenecks.
  • Design and execute experiments to validate performance hypotheses and benchmark new approaches.
  • Partner with Research Scientists to productionize novel training methods and contribute to publications and open-source releases.
  • Rapidly enable newly released open-source foundation models on the platform.
  • Build and maintain experimental infrastructure with production-quality reliability and scalability.

Requirements

  • Independently take ambiguous performance or infrastructure problems from investigation through deployment.
  • Strong programming skills in Python and PyTorch, with an emphasis on efficient, maintainable code.
  • Hands-on experience training or fine-tuning large neural networks in multi-GPU or multi-node environments.
  • Solid understanding of GPU architecture, mixed-precision training, and distributed training paradigms including data, tensor, pipeline, or expert parallelism.
  • Strong communication and collaboration skills across research and engineering teams.
  • Experience with optimized NVIDIA GPU kernels using CUDA or Triton, or communication collectives using NCCL or NVSHMEM, is a plus.
  • Experience with FSDP, DeepSpeed, Megatron-LM, or custom distributed training systems is a plus.
  • Experience optimizing distributed training for compute efficiency, memory efficiency, or scalability is a plus.
  • Experience managing large-scale GPU experiments, including scheduling, monitoring, and fault tolerance, is a plus.
  • Contributions to widely used open-source ML or ML systems projects are a plus.
  • Experience building or operating ML products or managed services for external customers is a plus.

Benefits

  • Competitive compensation, startup equity, health insurance, and other benefits are offered.
  • This is a full-time position with a stated US base salary range.

Tech Stack

Categories

AI ResearchML Engineering
Together AI

About Together AI

201-500 employees

Together AI builds an AI-native cloud platform for developers, offering high-performance inference, fine-tuning/model shaping, and large-scale pre-training on on-demand GPU clusters with APIs and managed services. It emphasizes open-source models that teams can run and adapt, and also provides infrastructure for decentralized and scalable workloads. Founded in 2022 and headquartered in San Francisco, it is privately held and reports notable customers including Cursor, ElevenLabs, Salesforce, and Zoom.

Contact me