Together AI

Research Engineer, Post-Training Inference

Together AI
Apply
2 months ago
San Francisco, CA, USAMid Level
H1B Sponsor

Base Salary

$200k - $290k/yr

Responsibilities

  • Design and build systems for customizing open-source models.
  • Integrate Model Shaping and Inference platforms to connect post-training with production serving.
  • Add inference-engine features and optimize reinforcement-learning workloads for large-scale post-training experiments.
  • Maintain a stable, robust platform through on-call participation and 24/7 availability support.
  • Collaborate with product, research, and engineering teams to keep APIs reliable, performant, and integrated with company infrastructure.

Requirements

  • At least 2 years of experience building and deploying machine-learning services in production.
  • Hands-on experience with modern inference engines such as SGLang, vLLM, and TensorRT-LLM.
  • Familiarity with current fine-tuning methods for LLMs and other AI models.
  • Strong software engineering background in Python or Go.
  • Experience with low-precision FP4/FP8 serving, Multi-LoRA, or models distributed across multiple GPU nodes is preferred.
  • Experience optimizing reinforcement-learning workloads is preferred.
  • Experience developing CUDA, Triton, or CuTE DSL inference kernels is preferred.
  • Experience developing large-scale, high-load production systems is preferred.
  • Experience maintaining or contributing to open-source ML projects is preferred.
  • Experience managing machine-learning workloads on Kubernetes clusters is preferred.

Benefits

  • Competitive compensation, startup equity, health insurance, and other benefits.
  • Full-time position; US base salary range is $200,000-$290,000.
Together AI

About Together AI

201-500 employees

Together AI is the AI Native Cloud, purpose-built for AI engineers and researchers with a full suite of tooling across inference, model shaping, and pre-training. AI natives can use Together AI as a full-stack AI platform — from a high- performance inference engine built for reliable and fast scaling to on-demand GPU clusters and massive-scale AI factories. Together AI continuously pushes the frontier forward by productizing cutting-edge research from our world-leading AI systems research team. By combining research velocity with production-grade infrastructure, we enable companies to reliably scale AI-native applications as fast as the field evolves. Trusted by leading AI natives like Cursor, Decagon, Eleven Labs, AI21, Hedra, and Cartesia, as well as SaaS innovators such as Salesforce, Zoom, and Zomato, Together AI powers the next generation of AI-native applications.