Plaud

Machine Learning Engineer, Inference & Serving (Speech LLM) - San Francisco

Plaud
Apply
4 months ago

Base Salary

$180k - $270k/yr

Responsibilities

  • Build and deploy high-throughput, ultra-low-latency inference engines for speech and large language models.
  • Optimize continuous batching, KV-cache management, stateful connections, and streaming inference latency.
  • Identify and eliminate GPU and memory-hierarchy bottlenecks across NVIDIA GPU clusters.
  • Develop real-time audio streaming and chunked generation systems for conversational AI.
  • Implement advanced decoding, model compression, quantization, and distributed inference techniques.
  • Collaborate with ML training and backend infrastructure teams to deliver speech-native AI systems.

Requirements

  • Hands-on experience building and deploying high-throughput, ultra-low-latency inference engines for large language models or foundational speech models.
  • Strong understanding of latency, throughput, and time-to-first-token or time-to-first-audio tradeoffs in real-time streaming environments.
  • Practical experience with continuous batching, KV-cache management such as PagedAttention, and stateful connections for real-time conversational AI.
  • Deep understanding of NVIDIA Ampere and Hopper GPU architectures and memory hierarchy.
  • Ability to collaborate across ML training and backend infrastructure teams in a fast-moving environment.
  • Preferred experience with vLLM, TensorRT-LLM, SGLang, or NVIDIA Triton Inference Server, including open-source contributions.
  • Preferred experience with WebSockets or WebRTC audio streaming, neural audio codecs, and chunked audio generation.
  • Preferred experience implementing speculative decoding, lookahead decoding, or chunked prefill.
  • Preferred experience with post-training quantization and deploying models using FP8, INT8, AWQ, or GPTQ.
  • Preferred experience deploying multi-GPU and multi-node inference pipelines and managing autoscaling infrastructure with Kubernetes.

Benefits

  • Top-tier healthcare, dental, and vision coverage for employees and dependents with an employer subsidy.
  • 401(k) plan with company matching for full-time employees.
  • Unlimited paid time off and 13 paid holidays.
  • 12 weeks of paid new-parent leave regardless of gender.
  • Hybrid work arrangement requiring at least three days in the office per week in San Francisco.
  • Choice of high-end laptops or workstations, annual offsites, and a fully stocked office.
Plaud

About Plaud

501-1,000 employees

Plaud is building the world's most trusted AI work companion for professionals to elevate productivity and performance through note-taking solutions, loved by over 2,000,000 users worldwide since 2023. With a mission to amplify human intelligence, Plaud is building the next-generation intelligence infrastructure and interfaces to capture, extract, and utilize what you say, hear, see, and think. Plaud Inc. is a Delaware-incorporated, San Francisco-based company pushing the boundary of human–AI intelligence through a hardware–software combination. With full ISO 27001, ISO 27701, GDPR, SOC 2, HIPAA, and EN18031 compliance, Plaud is committed to the highest standards of data security and privacy protection. To learn more about Plaud, please visit https://www.plaud.ai