Together AI

Staff Machine Learning Engineer, Voice AI

Together AI
Apply
3 months ago

Base Salary

$220k - $280k/yr

Responsibilities

  • Own the end-to-end voice inference roadmap and technical strategy across STT, TTS, and speech-to-speech models.
  • Architect and implement high-performance inference systems targeting leading time-to-first-byte, throughput, and GPU utilization.
  • Design production serving architectures for serverless and dedicated endpoints, including batching, streaming inference, memory management, reliability, and latency SLAs.
  • Build an extensible voice evaluation platform covering STT word error rate, accents, languages, and noise conditions, plus TTS naturalness, latency, and pronunciation fidelity.
  • Enable support for emerging audio-native LLM, codec-based, and end-to-end speech-to-speech architectures.
  • Lead model-partner integrations with Cartesia, Deepgram, Rime, and other partners from integration through optimization and ongoing performance accountability.
  • Diagnose GPU-kernel and framework bottlenecks through profiling and root-cause analysis, and deliver measurable performance improvements.
  • Influence platform architecture to meet the reliability and latency requirements of real-time voice APIs.
  • Define the technical direction and infrastructure primitives for customer fine-tuning of STT and TTS models.
  • Establish scalable technical foundations for multiple future voice products.

Requirements

  • 8+ years of ML engineering experience focused on model serving, inference optimization, or production-scale ML infrastructure.
  • Deep practical expertise in LLM serving engines such as vLLM, SGLang, TensorRT-LLM, or equivalent, including modifying internals and debugging production edge cases.
  • Expert proficiency in Python and PyTorch, with strong knowledge of CUDA kernels, GPU memory hierarchies, and profiling toolchains.
  • Proven system-design judgment and experience making architectural decisions that scale and influence platform evolution.
  • Strong technical leadership, autonomy, product intuition for developer tooling, and ability to work effectively in ambiguous early-stage environments.
  • Strong foundation in speech and audio ML, including ASR/TTS architectures and audio signal processing; directly relevant experience is preferred.
  • Familiarity with SNAC, Encodec, and DAC audio codec and tokenization schemes is a plus.
  • Experience training or fine-tuning speech models at scale is a significant advantage.
  • Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related field, or equivalent depth demonstrated through work.

Benefits

  • Full-time position with startup equity and health insurance.
  • Competitive benefits are offered; work arrangement is not specified.
Together AI

About Together AI

201-500 employees

Together AI is the AI Native Cloud, purpose-built for AI engineers and researchers with a full suite of tooling across inference, model shaping, and pre-training. AI natives can use Together AI as a full-stack AI platform — from a high- performance inference engine built for reliable and fast scaling to on-demand GPU clusters and massive-scale AI factories. Together AI continuously pushes the frontier forward by productizing cutting-edge research from our world-leading AI systems research team. By combining research velocity with production-grade infrastructure, we enable companies to reliably scale AI-native applications as fast as the field evolves. Trusted by leading AI natives like Cursor, Decagon, Eleven Labs, AI21, Hedra, and Cartesia, as well as SaaS innovators such as Salesforce, Zoom, and Zomato, Together AI powers the next generation of AI-native applications.