about 3 hours ago
Remote, United StatesMid Level / Senior
H1B Sponsor
Base Salary
$200k - $220k/yr
Responsibilities
- Architect, implement, pre-train, fine-tune, and post-train large-scale speech models.
- Design, run, and analyze scientific experiments to enhance model understanding.
- Develop and improve development tooling to boost team productivity.
- Contribute to the entire stack from low-level optimizations to high-level model design.
- Define data requirements and collaborate on data acquisition and quality.
- Design automated evaluations and conduct robustness and bias checks.
- Harden the training, evaluation, and inference pipeline to meet production SLAs.
- Contribute to safety and consent guardrails for responsible speech technology.
Requirements
- Exceptional experience with large-scale audio models over 8B parameters.
- Hands-on experience with diffusion and flow-matching transformers.
- Experience training audio VAEs, neural audio codecs, and vocoders.
- Strong experience with multi-node, multi-GPU distributed training.
- Proven software engineering skills in building complex systems.
- Proficiency in PyTorch and performance optimization techniques.
- Experience shipping large-scale speech or multimodal generative models.
- Background in working with large-scale ML data and quality triangulation.
- Experience with voice cloning and expressive speech generation.
- Notable publications or open-source contributions in the field.
Benefits
- Competitive salary and generous company equity.
- 99.99% of medical, dental, and vision insurance premiums covered.
- 42 days of paid time off, including PTO, sick days, and holidays.
- Generous parental leave and fertility support.
- 401(k) retirement savings plan.
- Lifestyle spending account of $500/month.
- Complimentary lunch and snacks for in-office employees.
- One Medical membership and more.
