3 days ago
San Jose, CA, USAStaff+
Base Salary
$300k - $500k/yr
Responsibilities
- Evaluate GPUs, NPUs, DSPs, and specialized accelerators for on-device deployment and own hardware recommendation decisions.
- Co-design model architectures with foundation model and audio ML teams to meet latency, memory, and power constraints.
- Build low-level execution layers, custom kernels, runtime systems, and compiler paths for transformer workloads on target hardware.
- Partner with silicon vendors and internal hardware teams to bring up accelerators and optimize transformer execution.
- Hire and lead engineers working on performance-critical software and set the technical direction for the inference stack.
Requirements
- Require 8–12+ years in high-performance computing, including production workloads deployed on GPUs, NPUs, or specialized accelerators.
- Require deep knowledge of attention, KV-cache behavior, quantization effects, and memory bandwidth limitations.
- Require experience designing or optimizing inference engines, distributed runtimes, or ML compilers, including writing kernels when necessary.
- Require experience leading teams developing performance-critical software and setting technical direction for a software stack.
- Require experience taking a model from a research checkpoint to constrained hardware in a product used by customers.
- Preferred experience includes Hexagon DSP, Ambiq-class MCUs, or comparable embedded AI silicon.
- Preferred experience includes speech, audio, or streaming multimodal inference.
- Preferred contributions to open-source inference or compiler toolchains such as TensorRT, ONNX Runtime, TVM, or MLIR.
Benefits
- Full-time position with a US base salary range of $300,000–$500,000 annually.
- Total compensation may include additional components and benefits, with details shared if an employment offer is extended.
