7 months ago
Responsibilities
- Implement and optimize inference kernels for CPU, NPU, and GPU architectures across edge hardware
- Develop quantization strategies using INT4, INT8, and FP8 while preserving model quality within strict memory budgets
- Contribute to llama.cpp and other open-source inference frameworks, including support for audio and vision model architectures
- Profile and optimize inference pipelines to achieve sub-100ms time-to-first-token performance on target devices
- Collaborate with ML researchers to identify optimization opportunities for Liquid Foundation Models
- Own major workstreams and ship measurable latency or memory improvements to production devices
Requirements
- 5+ years of systems programming experience with strong C++ proficiency
- Experience in embedded software engineering or working with resource-constrained systems
- Understanding of ML fundamentals at the linear algebra level, including matrix operations, attention, and quantization
- Understanding of hardware architecture concepts including cache hierarchies, memory bandwidth, and SIMD/vectorization
- Experience contributing to llama.cpp, ExecuTorch, or similar inference frameworks is preferred
- Rust systems programming experience is preferred
- Background in custom accelerator development or accelerator teams is preferred
- A quantitative degree in mathematics, physics, or a similar field combined with engineering experience is preferred
Benefits
- 100% of medical, dental, and vision premiums covered for employees and dependents
- 401(k) matching available
- Unlimited PTO and company-wide Refill Days
- Open to locations beyond the preferred San Francisco and Boston offices
- Opportunity to work on novel AI optimization challenges with code deployed to real devices
Categories
About Liquid AI
We build efficient general-purpose AI at every scale.
