Centific

AI Research Engineer- Speech 1

Centific
Apply
2 months ago
Redmond, WA, USAMid Level

Base Salary

$150k - $160k/yr

Responsibilities

  • Design, develop, and deploy Large Audio Language Models for native audio understanding, reasoning, and generation.
  • Build Large Audio Reasoning Models for complex reasoning over speech and audio inputs across medical, technical, and conversational domains.
  • Develop Speech-to-Speech systems covering speech understanding, dialogue management, and speech synthesis.
  • Research alignment between speech encoders and LLM backbones using adapters, LoRA, and efficient fine-tuning.
  • Design speech tokenization and temporal compression methods for long-form audio reasoning and multi-turn dialogue.
  • Build evaluation frameworks and benchmarks for speech question answering, audio understanding, and reasoning accuracy.
  • Optimize inference pipelines for low-latency and streaming speech applications.
  • Transfer research innovations into production systems and customer-facing applications.
  • Contribute to technical documentation, research write-ups, and publications at venues including NeurIPS, ICML, ACL, and Interspeech.

Requirements

  • Master's degree in Computer Science, Electrical Engineering, or a related field with a focus on speech, audio machine learning, or multimodal learning; Ph.D. preferred.
  • At least 2 years of industry or applied research experience in speech/audio AI, Large Language Models, or multimodal systems.
  • Applied research contributions demonstrated through publications, patents, or shipped speech/audio AI or LLM products.
  • Strong proficiency in Python and PyTorch, including hands-on GPU-accelerated training of large-scale models.
  • Understanding of speech and audio signal processing, acoustic modeling, and audio representations.
  • Working knowledge of Transformers, SSMs, instruction tuning, and alignment methods.
  • Familiarity with adapter-based integration, cross-modal attention, and audio-text fusion.
  • Preferred experience with Large Audio Language Models such as Qwen-Audio, SALMONN, LTU, or Gemini Audio.
  • Preferred experience with HuBERT, Wav2Vec 2.0, Whisper, WavLM, neural audio codecs, vocoders, speech synthesis, and discrete speech tokenization.
  • Preferred experience with audio reasoning benchmarks, distributed training, inference optimization, speech frameworks, multilingual speech, domain adaptation, or audio-model safety evaluation.
  • Preferred publication record at top-tier venues including NeurIPS, ICML, ICLR, ACL, Interspeech, or ICASSP.

Benefits

  • Competitive compensation package with comprehensive benefits.
  • Opportunity to work on Large Audio Language Models and audio reasoning research with real-world impact.
  • Collaboration with experienced applied scientists and engineers in speech and multimodal AI.
  • Support for publications at top-tier conferences and professional development.
  • Access to state-of-the-art GPU infrastructure for large-scale audio-model training.
  • Flexible hybrid or remote work arrangements; locations include Redmond, Washington, and Palo Alto, California.

Tech Stack

FastAPIgRPCHugging Face TransformersPythonPyTorch

Categories

AI Research
Centific

About Centific

1,001-5,000 employees
Contact me