Cantina

ML Research Engineer, TTS

Cantina
Apply
4 months ago
Barcelona, SpainSenior
H1B Sponsor

Responsibilities

  • Architect, implement, pre-train, fine-tune, and post-train or align large-scale speech models.
  • Independently lead small research projects and collaborate on larger team initiatives.
  • Design, run, and analyze scientific experiments.
  • Develop tooling that improves research and engineering productivity.
  • Contribute across the stack, from low-level optimizations to high-level model design.
  • Define data requirements and partner on data acquisition, curation, augmentation, labeling quality, and synthetic-data strategies.
  • Design automated objective and subjective evaluations, including listening tests, SV/WER/ASR metrics, robustness and bias checks, and red-team studies.
  • Harden training, evaluation, and inference pipelines; profile latency, memory, and cost; and meet production SLAs with monitoring and rollback.
  • Partner with infrastructure teams on distributed training and inference across cloud GPU fleets and productionize models with reliability and observability.
  • Contribute to safety and consent guardrails and misuse or abuse mitigation for speech technology.

Requirements

  • Exceptional research and development experience with large-scale audio models exceeding 3 billion parameters and more than 500,000 hours of data.
  • Exceptional understanding and hands-on experience with transformer architectures, diffusion models including distillation and streaming, and/or audio language modeling.
  • Strong experience with multi-node and multi-GPU distributed model training.
  • Strong software engineering skills and a track record of building complex systems.
  • Strong PyTorch and performance optimization skills, including profiling and use of CUDA, Triton, or C++ as needed.
  • Experience shipping large-scale speech or audio models to production.
  • Background working with large-scale machine-learning data.
  • Ability to iterate on data and assess quality using subjective and objective signals.
  • Notable publications and/or open-source contributions in speech, audio, or machine learning.
  • Experience with voice cloning, speech control, or voice generation.
  • Preferred experience includes production work with large-scale TTS, voice conversion, or ASR models; large-scale ML systems; audio language modeling; transformer architectures; large-scale ML data processing; and speech, audio, or ML publications or open-source work.

Tech Stack

Categories

AI ResearchML Engineering
Cantina

About Cantina

201-500 employees

Cantina Labs is a social AI company, developing a suite of advanced real-time models that push the boundaries of expression, personality, and realism. We bring characters to life, transforming how people tell stories, connect, and create. We build and power ecosystems. Cantina, our flagship social AI platform, is just the beginning.