Deepgram

Embedded AI Engineer, On-Device Models

Deepgram
Apply
2 months ago
Remote, Worldwide or New York, NY, USASenior / Staff+
H1B Sponsor

Base Salary

$219k - $274k/yr

Responsibilities

  • Deploy Deepgram speech and conversational models on embedded and low-power consumer hardware across diverse processors and accelerators.
  • Optimize models using quantization, pruning, distillation, operator fusion, and architecture-specific compilation for latency, memory, power, and thermal constraints.
  • Write and optimize performance-critical C, C++, and/or Rust runtime code for embedded, bare-metal, and real-time operating system environments.
  • Integrate edge inference runtimes and vendor NPU/DSP toolchains across CPU, GPU, NPU, and accelerator architectures.
  • Build model packaging, deployment pipelines, over-the-air updates, and lightweight telemetry for intermittently connected devices.
  • Establish benchmarking and validation for latency, accuracy, power consumption, memory footprint, and resource utilization.
  • Partner with silicon and device vendors on SDK integration, performance tuning, and new chipset and reference-platform support.
  • Collaborate with Research and Engine teams to shape models for efficient edge deployment.

Requirements

  • Experience delivering production systems on resource-constrained hardware such as embedded systems, mobile, edge AI, or low-power devices.
  • Strong proficiency in C, C++, and/or Rust with experience writing performance-critical constrained-environment code.
  • Hands-on experience with on-device model optimization, including quantization, pruning, knowledge distillation, or architecture-specific compilation.
  • Familiarity with edge inference runtimes such as ONNX Runtime, TensorRT, TFLite, or ExecuTorch, and/or vendor NPU/DSP toolchains.
  • Understanding of CPU, GPU, NPU, and DSP architectures, memory hierarchies, fixed-point and integer arithmetic, and power management.
  • Experience with bare-metal or RTOS environments such as FreeRTOS or Zephyr, embedded Linux, microcontrollers, or edge SoCs.
  • Strong communication skills and the ability to scope ambiguous optimization problems, deliver measurable results, and explain tradeoffs.
  • Preferred: experience with real-time embedded audio processing, DSP pipelines, audio codec optimization, wake-word or always-on listening, or streaming inference.
  • Preferred: advanced ML optimization including custom quantization, mixed-precision inference, or neural architecture search.
  • Preferred: hardware evaluation and benchmarking across accelerators, SoCs, or GPUs.
  • Preferred: experience shipping AI features in scaled consumer products.
  • Preferred: familiarity with model compilation and optimization toolchains across hardware targets.
  • Preferred: experience with secure on-device deployment, code signing, encrypted model storage, and safe update mechanisms.

Tech Stack

Deepgram

About Deepgram

51-200 employees

Deepgram is the real-time API platform powering the trillion-dollar Voice AI economy. Backed by a $130M Series C at a $1.3B valuation, Deepgram is trusted by 200,000+ developers and 1,300+ organizations to build Voice AI products, platforms, and autonomous agents with the lowest latency, highest accuracy, and enterprise reliability. Our voice-native foundation models and runtime infrastructure have processed 50,000+ years of audio and over 1 trillion words, making Deepgram the most experienced voice AI platform in the world. Industry-leading models & platform: 👂 Nova-3 — the world’s most accurate real-time speech-to-text model 🔊 Aura-2 — professional, enterprise-grade text-to-speech 💬 Flux — the first Conversational Speech Recognition model designed to handle interruptions 🚀 Voice Agent API — enterprise-ready, real-time conversational AI 🧠 Saga — the Voice OS Beyond core infrastructure, Deepgram is expanding the Voice AI ecosystem through: 💪 Powered by Deepgram, supporting voice products built by leading AI startups and enterprise organizations 🌉 A new Voice AI Collaboration Hub in San Francisco for builders, partners, and the voice community 🍔 The acquisition of OfOne, delivering real-time Voice AI for restaurants and drive-thru operations with 95%+ containment 📃 A growing patent portfolio in Voice AI Much like APIs powered the payments and cloud economies, Deepgram is building the foundation for a trillion-dollar B2B Voice AI economy—centered on the most natural human interface: voice.