GrepJob
Cantina

Machine Learning Engineer - Voice Conversion

Cantina
Apply
about 3 hours ago
Remote, United StatesMid Level / Senior
H1B Sponsor

Base Salary

$200k - $220k/yr

Responsibilities

  • Architect, implement, pre-train, fine-tune, and post-train large-scale speech models.
  • Design, run, and analyze scientific experiments to enhance model understanding.
  • Develop and improve development tooling to boost team productivity.
  • Contribute to the entire stack from low-level optimizations to high-level model design.
  • Define data requirements and collaborate on data acquisition and quality.
  • Design automated evaluations and conduct robustness and bias checks.
  • Harden the training, evaluation, and inference pipeline to meet production SLAs.
  • Contribute to safety and consent guardrails for responsible speech technology.

Requirements

  • Exceptional experience with large-scale audio models over 8B parameters.
  • Hands-on experience with diffusion and flow-matching transformers.
  • Experience training audio VAEs, neural audio codecs, and vocoders.
  • Strong experience with multi-node, multi-GPU distributed training.
  • Proven software engineering skills in building complex systems.
  • Proficiency in PyTorch and performance optimization techniques.
  • Experience shipping large-scale speech or multimodal generative models.
  • Background in working with large-scale ML data and quality triangulation.
  • Experience with voice cloning and expressive speech generation.
  • Notable publications or open-source contributions in the field.

Benefits

  • Competitive salary and generous company equity.
  • 99.99% of medical, dental, and vision insurance premiums covered.
  • 42 days of paid time off, including PTO, sick days, and holidays.
  • Generous parental leave and fertility support.
  • 401(k) retirement savings plan.
  • Lifestyle spending account of $500/month.
  • Complimentary lunch and snacks for in-office employees.
  • One Medical membership and more.

Tech Stack

Categories