Figure AI

Staff AI Inference and Acceleration Engineer

Figure AI
Apply
1 month ago
San Jose, CA, USAStaff+

Base Salary

$180k - $275k/yr

Responsibilities

  • Own the on-board inference architecture and map models to available NPU, GPU, DSP, and CPU resources based on latency, power, and memory budgets.
  • Partition inference workloads across heterogeneous compute resources while balancing real-time performance, power, and thermal constraints.
  • Define and maintain a system-level compute budget for all inference tasks running on the robot.
  • Evaluate next-generation acceleration hardware and help define future compute platform requirements.
  • Optimize inference toolchains from model export through runtime execution for target hardware.
  • Apply quantization, pruning, operator fusion, and other compression techniques to reduce compute, memory, and power usage.
  • Profile inference pipelines and optimize kernel scheduling, memory layout, and data movement.
  • Collaborate with AI/ML and Platform Software teams on hardware-friendly model architectures, runtime integration, scheduling, and power management.
  • Engage with silicon vendors and research teams to track accelerator developments and influence hardware roadmaps.

Requirements

  • M.S. or Ph.D. in Computer Engineering, Electrical Engineering, Computer Science, or a related field, or equivalent industry experience.
  • At least 8 years of industry experience in hardware acceleration, ML systems, or compute architecture.
  • Deep understanding of AI/ML inference, model formats, inference runtimes, and deployment pipelines.
  • Hands-on experience optimizing models for edge or embedded hardware using quantization, pruning, and operator-level tuning.
  • Strong understanding of computer architecture, including memory hierarchies, data movement, and heterogeneous compute.
  • Experience profiling and benchmarking inference workloads across CPU, GPU, NPU, and DSP.
  • Familiarity with low-level toolchains and compilation frameworks including TVM, MLIR, TensorRT, Torch, SNPE/QNN, JAX, CUDA, and ROCm.
  • Strong software engineering skills in C++ and Python.
  • Ability to work effectively across hardware, software, and AI/ML teams.
  • Knowledge of real-time operating constraints and experience co-designing model architectures with ML teams are bonus qualifications.
Contact me