Together AI

Forward Deployed Engineer (Inference & Post-Training)

Together AI
Apply
19 hours ago

Base Salary

$270k - $300k/yr

Responsibilities

  • Select, configure, and optimize inference engines based on hardware, model architecture, and workload profile.
  • Develop configuration updates and tune KV cache, speculative decoding, tensor parallelism, and quantization to meet throughput and latency targets.
  • Run and optimize reinforcement-learning training and guide customers through LoRA, SFT, DPO, RLHF, and GRPO pipelines.
  • Serve as the primary technical contact for strategic accounts, optimizing endpoint configurations and supporting customer milestones.
  • Establish opinionated onboarding configurations and improve customer time-to-value.
  • Surface field insights, contribute to software and model roadmap improvements, and drive early feature and research adoption.

Requirements

  • 5+ years in a technical role focused on inference systems, open-source LLM deployment, or post-training workflows.
  • Expert hands-on experience with inference engines such as vLLM, TensorRT-LLM, or SGLang, including diagnosing engine-level performance issues.
  • Deep knowledge of KV cache tuning, speculative decoding, tensor parallelism, pipeline parallelism, and quantization.
  • Hands-on experience with fine-tuning and post-training pipelines including LoRA, SFT, DPO, RLHF, and GRPO.
  • Broad knowledge of open-source models and judgment in selecting models for customer use cases, hardware profiles, and performance targets.
  • Strong Python skills and comfort working in production environments.

Benefits

  • US base salary range of $270,000-$300,000 per year, plus startup equity and benefits.
  • Health insurance and other benefits are offered.
  • Flexible remote-work arrangements are available.
  • This is a full-time position.

Tech Stack

Categories

Forward Deployed
Together AI

About Together AI

201-500 employees

Together AI builds an AI-native cloud platform for developers, offering high-performance inference, fine-tuning/model shaping, and large-scale pre-training on on-demand GPU clusters with APIs and managed services. It emphasizes open-source models that teams can run and adapt, and also provides infrastructure for decentralized and scalable workloads. Founded in 2022 and headquartered in San Francisco, it is privately held and reports notable customers including Cursor, ElevenLabs, Salesforce, and Zoom.

Contact me