2 months ago
Redmond, WA, USAMid Level
Base Salary
$150k - $160k/yr
Responsibilities
- Design, develop, and deploy Large Audio Language Models for native audio understanding, reasoning, and generation.
- Build Large Audio Reasoning Models for complex reasoning over speech and audio inputs across medical, technical, and conversational domains.
- Develop Speech-to-Speech systems covering speech understanding, dialogue management, and speech synthesis.
- Research alignment between speech encoders and LLM backbones using adapters, LoRA, and efficient fine-tuning.
- Design speech tokenization and temporal compression methods for long-form audio reasoning and multi-turn dialogue.
- Build evaluation frameworks and benchmarks for speech question answering, audio understanding, and reasoning accuracy.
- Optimize inference pipelines for low-latency and streaming speech applications.
- Transfer research innovations into production systems and customer-facing applications.
- Contribute to technical documentation, research write-ups, and publications at venues including NeurIPS, ICML, ACL, and Interspeech.
Requirements
- Master's degree in Computer Science, Electrical Engineering, or a related field with a focus on speech, audio machine learning, or multimodal learning; Ph.D. preferred.
- At least 2 years of industry or applied research experience in speech/audio AI, Large Language Models, or multimodal systems.
- Applied research contributions demonstrated through publications, patents, or shipped speech/audio AI or LLM products.
- Strong proficiency in Python and PyTorch, including hands-on GPU-accelerated training of large-scale models.
- Understanding of speech and audio signal processing, acoustic modeling, and audio representations.
- Working knowledge of Transformers, SSMs, instruction tuning, and alignment methods.
- Familiarity with adapter-based integration, cross-modal attention, and audio-text fusion.
- Preferred experience with Large Audio Language Models such as Qwen-Audio, SALMONN, LTU, or Gemini Audio.
- Preferred experience with HuBERT, Wav2Vec 2.0, Whisper, WavLM, neural audio codecs, vocoders, speech synthesis, and discrete speech tokenization.
- Preferred experience with audio reasoning benchmarks, distributed training, inference optimization, speech frameworks, multilingual speech, domain adaptation, or audio-model safety evaluation.
- Preferred publication record at top-tier venues including NeurIPS, ICML, ICLR, ACL, Interspeech, or ICASSP.
Benefits
- Competitive compensation package with comprehensive benefits.
- Opportunity to work on Large Audio Language Models and audio reasoning research with real-world impact.
- Collaboration with experienced applied scientists and engineers in speech and multimodal AI.
- Support for publications at top-tier conferences and professional development.
- Access to state-of-the-art GPU infrastructure for large-scale audio-model training.
- Flexible hybrid or remote work arrangements; locations include Redmond, Washington, and Palo Alto, California.
