3 months ago
Remote, United States or New York, NY, USASenior / Staff+
H1B sponsor
Base Salary
$219k - $274k/yr
Responsibilities
- Deploy Deepgram speech and conversational models on embedded and low-power consumer hardware across diverse processors and accelerators.
- Optimize models using quantization, pruning, distillation, operator fusion, and architecture-specific compilation for latency, memory, power, and thermal constraints.
- Write and optimize performance-critical C, C++, and/or Rust runtime code for embedded, bare-metal, and real-time operating system environments.
- Integrate edge inference runtimes and vendor NPU/DSP toolchains across CPU, GPU, NPU, and accelerator architectures.
- Build model packaging, deployment pipelines, over-the-air updates, and lightweight telemetry for intermittently connected devices.
- Establish benchmarking and validation for latency, accuracy, power consumption, memory footprint, and resource utilization.
- Partner with silicon and device vendors on SDK integration, performance tuning, and new chipset and reference-platform support.
- Collaborate with Research and Engine teams to shape models for efficient edge deployment.
Requirements
- Experience delivering production systems on resource-constrained hardware such as embedded systems, mobile, edge AI, or low-power devices.
- Strong proficiency in C, C++, and/or Rust with experience writing performance-critical constrained-environment code.
- Hands-on experience with on-device model optimization, including quantization, pruning, knowledge distillation, or architecture-specific compilation.
- Familiarity with edge inference runtimes such as ONNX Runtime, TensorRT, TFLite, or ExecuTorch, and/or vendor NPU/DSP toolchains.
- Understanding of CPU, GPU, NPU, and DSP architectures, memory hierarchies, fixed-point and integer arithmetic, and power management.
- Experience with bare-metal or RTOS environments such as FreeRTOS or Zephyr, embedded Linux, microcontrollers, or edge SoCs.
- Strong communication skills and the ability to scope ambiguous optimization problems, deliver measurable results, and explain tradeoffs.
- Preferred: experience with real-time embedded audio processing, DSP pipelines, audio codec optimization, wake-word or always-on listening, or streaming inference.
- Preferred: advanced ML optimization including custom quantization, mixed-precision inference, or neural architecture search.
- Preferred: hardware evaluation and benchmarking across accelerators, SoCs, or GPUs.
- Preferred: experience shipping AI features in scaled consumer products.
- Preferred: familiarity with model compilation and optimization toolchains across hardware targets.
- Preferred: experience with secure on-device deployment, code signing, encrypted model storage, and safe update mechanisms.
Categories
About Deepgram
Deepgram builds Voice AI infrastructure for developers, offering speech-to-text, text-to-speech, and voice agent tooling through real-time APIs and deployable on-prem software. Privately held and founded in 2015, the company is headquartered in San Francisco. Customers such as Twilio, Cloudflare, and Jack in the Box use its models to power production voice features; products include Nova-3 STT, Aura TTS, and a Voice Agent API.
