Liquid AI

Member of Technical Staff - Edge Inference Engineer

Liquid AI
Apply
7 months ago
Remote, Worldwide +2 moreSenior
H1B Sponsor

Responsibilities

  • Implement and optimize inference kernels for CPU, NPU, and GPU architectures across edge hardware
  • Develop quantization strategies using INT4, INT8, and FP8 while preserving model quality within strict memory budgets
  • Contribute to llama.cpp and other open-source inference frameworks, including support for audio and vision model architectures
  • Profile and optimize inference pipelines to achieve sub-100ms time-to-first-token performance on target devices
  • Collaborate with ML researchers to identify optimization opportunities for Liquid Foundation Models
  • Own major workstreams and ship measurable latency or memory improvements to production devices

Requirements

  • 5+ years of systems programming experience with strong C++ proficiency
  • Experience in embedded software engineering or working with resource-constrained systems
  • Understanding of ML fundamentals at the linear algebra level, including matrix operations, attention, and quantization
  • Understanding of hardware architecture concepts including cache hierarchies, memory bandwidth, and SIMD/vectorization
  • Experience contributing to llama.cpp, ExecuTorch, or similar inference frameworks is preferred
  • Rust systems programming experience is preferred
  • Background in custom accelerator development or accelerator teams is preferred
  • A quantitative degree in mathematics, physics, or a similar field combined with engineering experience is preferred

Benefits

  • 100% of medical, dental, and vision premiums covered for employees and dependents
  • 401(k) matching available
  • Unlimited PTO and company-wide Refill Days
  • Open to locations beyond the preferred San Francisco and Boston offices
  • Opportunity to work on novel AI optimization challenges with code deployed to real devices

Tech Stack

Liquid AI

About Liquid AI

51-200 employees

We build efficient general-purpose AI at every scale.