Hark

On-Device AI Inference Engineer

Hark
Apply
3 days ago
San Jose, CA, USAMid Level / Senior

Base Salary

$200k - $450k/yr

Responsibilities

  • Write and optimize low-level kernels and runtime paths for transformer workloads on target silicon.
  • Design model memory residency, scheduling, and swap behavior across concurrent workloads with limited memory and power.
  • Profile models on real hardware, identify bottlenecks, and improve delivered performance.
  • Quantize models from full precision to INT8/INT4 and meet product size, latency, and power budgets.
  • Optimize transformer workloads for new accelerators in collaboration with the hardware team.
  • Provide deployment constraints and performance feedback to model teams.

Requirements

  • 4–8+ years of experience writing performance-critical software and optimizing workloads on GPUs, NPUs, DSPs, or similar accelerators.
  • Strong C/C++ skills and experience with SIMD, custom kernels, memory layout, and profiling tools.
  • Understanding of attention, KV-cache behavior, transformer inference performance, and memory bandwidth.
  • Experience optimizing against compute, memory, and power budgets.
  • Experience deploying an optimized model in a product on constrained hardware.
  • Preferred: experience with Hexagon DSP, Ambiq-class MCUs, or comparable embedded AI silicon.
  • Preferred: familiarity with ONNX Runtime, TVM, MLIR, TensorRT, or similar inference and compiler toolchains.
  • Preferred: background in speech or audio inference.

Benefits

  • Full-time position with a stated US base salary range of $200,000–$450,000 annually.
  • Total compensation may include additional components and benefits depending on the specific role.

Tech Stack

Hark

About Hark

1-10 employees
Contact me