2 months ago
Base Salary
$200k - $220k/yr
Responsibilities
- Architect, implement, pre-train, fine-tune, and post-train or align large-scale speech models
- Design, run, and analyze scientific experiments to advance model understanding and quality
- Develop tooling that improves team productivity
- Contribute across the stack from low-level optimization to high-level model design
- Define data requirements and partner on acquisition, curation, augmentation, labeling quality, and synthetic data strategies
- Design automated objective and subjective evaluations, including listening tests, SV/WER/ASR-based metrics, robustness and bias checks, and red-team studies
- Harden the training-to-evaluation-to-inference pipeline and profile latency, memory, and cost
- Meet production SLAs through robust monitoring and rollback
- Contribute to safety and consent guardrails and misuse or abuse mitigation for speech technology
Requirements
- Exceptional research or development experience with large-scale audio models exceeding 8B parameters and 500,000 hours of data
- Deep hands-on experience with diffusion and/or flow-matching transformers, including samplers, schedules, conditioning, and distillation
- Deep hands-on experience training audio VAEs, neural audio codecs, and vocoders, including latent/tokenizer design, reconstruction and perceptual objectives, and adversarial training
- Strong experience with multi-node, multi-GPU distributed training using FSDP, DeepSpeed, or equivalent
- Strong software engineering skills and a record of building complex systems
- Strong PyTorch and performance-optimization skills, including profiling and use of CUDA, Triton, or C++ as needed
- Experience shipping large-scale speech, audio, or multimodal generative models to production
- Background working with large-scale ML data and evaluating quality through subjective and objective signals
- Experience with voice cloning, speech control or steerability, or expressive speech generation
- Notable publications and/or open-source contributions in speech, audio, or machine learning
Benefits
- Competitive salary and generous company equity
- Medical, dental, and vision insurance with 99.99% of premiums covered by Cantina
- 42 days of paid time off, including PTO, sick days, company holidays, and floating holidays
- Generous parental leave and fertility support
- 401(k) retirement savings plan
- $500-per-month lifestyle spending account
- Complimentary lunch and snacks for in-office employees
- One Medical membership
Categories
AI ResearchML Engineering
About Cantina
Cantina Labs is a social AI company, developing a suite of advanced real-time models that push the boundaries of expression, personality, and realism. We bring characters to life, transforming how people tell stories, connect, and create. We build and power ecosystems. Cantina, our flagship social AI platform, is just the beginning.
