Hark

Member of Technical Staff, Multimodal Speech

Hark
Apply
10 days ago
San Jose, CA, USASenior

Base Salary

$180k - $450k/yr

Responsibilities

  • Advance speech and audio capabilities in multimodal models, including speech recognition, synthesis, and understanding.
  • Develop large-scale speech and audio data pipelines for collection, filtering, alignment, and synthetic data generation.
  • Design and implement multimodal speech and audio models and real-time systems.
  • Build evaluation frameworks and benchmarks for speech quality, latency, robustness, and user experience.
  • Optimize models and systems for real-time performance, scalability, and production deployment.
  • Collaborate with product and engineering teams to translate research innovations into user-facing AI experiences.

Requirements

  • Proven track record of advancing speech or audio models through data, modeling, or training innovations.
  • Strong experience with speech/audio domains such as ASR, TTS, speech-to-speech, or audio foundation models.
  • Experience with large-scale machine learning systems and distributed training.
  • Strong background in data-driven experimentation, systematic evaluation, and model iteration.
  • Ability to drive end-to-end impact from research through production.

Categories

AI ResearchML Engineering
Hark

About Hark

1-10 employees
Contact me