Together AI

Research Engineer, Large-Scale Training

Together AI
Apply
14 days ago

Base Salary

$200k - $290k/yr

Responsibilities

  • Design, implement, and optimize core components of large-scale training infrastructure.
  • Integrate new model architectures, validate training correctness and convergence, and optimize production fine-tuning workloads.
  • Profile distributed training workloads and remove compute, memory, and communication bottlenecks.
  • Design and execute experiments to validate performance hypotheses and benchmark new approaches.
  • Partner with Research Scientists to productionize novel training methods and contribute to publications and open-source releases.
  • Rapidly enable newly released open-source foundation models on the platform.
  • Build and maintain experimental infrastructure with production-quality reliability and scalability.

Requirements

  • Independently take ambiguous performance or infrastructure problems from investigation through deployment.
  • Strong programming skills in Python and PyTorch, with an emphasis on efficient, maintainable code.
  • Hands-on experience training or fine-tuning large neural networks in multi-GPU or multi-node environments.
  • Solid understanding of GPU architecture, mixed-precision training, and distributed training paradigms including data, tensor, pipeline, or expert parallelism.
  • Strong communication and collaboration skills across research and engineering teams.
  • Experience with optimized NVIDIA GPU kernels using CUDA or Triton, or communication collectives using NCCL or NVSHMEM, is a plus.
  • Experience with FSDP, DeepSpeed, Megatron-LM, or custom distributed training systems is a plus.
  • Experience optimizing distributed training for compute efficiency, memory efficiency, or scalability is a plus.
  • Experience managing large-scale GPU experiments, including scheduling, monitoring, and fault tolerance, is a plus.
  • Contributions to widely used open-source ML or ML systems projects are a plus.
  • Experience building or operating ML products or managed services for external customers is a plus.

Benefits

  • Competitive compensation, startup equity, health insurance, and other benefits are offered.
  • This is a full-time position with a stated US base salary range.

Tech Stack

Categories

AI ResearchML Engineering
Together AI

About Together AI

201-500 employees

Together AI is the AI Native Cloud, purpose-built for AI engineers and researchers with a full suite of tooling across inference, model shaping, and pre-training. AI natives can use Together AI as a full-stack AI platform — from a high- performance inference engine built for reliable and fast scaling to on-demand GPU clusters and massive-scale AI factories. Together AI continuously pushes the frontier forward by productizing cutting-edge research from our world-leading AI systems research team. By combining research velocity with production-grade infrastructure, we enable companies to reliably scale AI-native applications as fast as the field evolves. Trusted by leading AI natives like Cursor, Decagon, Eleven Labs, AI21, Hedra, and Cartesia, as well as SaaS innovators such as Salesforce, Zoom, and Zomato, Together AI powers the next generation of AI-native applications.