3 days ago
San Jose, CA, USAMid Level / Senior
Base Salary
$200k - $450k/yr
Responsibilities
- Write and optimize low-level kernels and runtime paths for transformer workloads on target silicon.
- Design model memory residency, scheduling, and swap behavior across concurrent workloads with limited memory and power.
- Profile models on real hardware, identify bottlenecks, and improve delivered performance.
- Quantize models from full precision to INT8/INT4 and meet product size, latency, and power budgets.
- Optimize transformer workloads for new accelerators in collaboration with the hardware team.
- Provide deployment constraints and performance feedback to model teams.
Requirements
- 4–8+ years of experience writing performance-critical software and optimizing workloads on GPUs, NPUs, DSPs, or similar accelerators.
- Strong C/C++ skills and experience with SIMD, custom kernels, memory layout, and profiling tools.
- Understanding of attention, KV-cache behavior, transformer inference performance, and memory bandwidth.
- Experience optimizing against compute, memory, and power budgets.
- Experience deploying an optimized model in a product on constrained hardware.
- Preferred: experience with Hexagon DSP, Ambiq-class MCUs, or comparable embedded AI silicon.
- Preferred: familiarity with ONNX Runtime, TVM, MLIR, TensorRT, or similar inference and compiler toolchains.
- Preferred: background in speech or audio inference.
Benefits
- Full-time position with a stated US base salary range of $200,000–$450,000 annually.
- Total compensation may include additional components and benefits depending on the specific role.
Tech Stack
CC++
