7 hours ago
Responsibilities
- Design and build benchmark suites covering inference performance, model quality, and knowledge evaluation across hardware targets.
- Run external partner verifications, compare solutions against benchmarks, identify gaps, and clearly communicate findings.
- Port models such as LFM2 across runtimes and frameworks and verify correctness end-to-end.
- Maintain and extend the inference engine layer using llama.cpp, ONNX, and MLX as new model architectures emerge.
- Make benchmark results explainable, verifiable, reproducible, and independently trustworthy for internal teams and partners.
Requirements
- Hands-on experience with at least one inference framework such as llama.cpp, ONNX Runtime, or MLX, including internals and modification beyond basic usage.
- Experience designing and building benchmarking pipelines with methodology, validation, and reproducibility.
- Strong C++ and Python experience in performance-sensitive contexts.
- Solid understanding of inference fundamentals, including quantization, decoding strategies, and memory layout and their interactions.
- Preferred experience porting models across runtimes and verifying numerical correctness.
- Preferred experience working with external partners or clients in technical validation or evaluation.
- Preferred familiarity with edge inference targets and their constraints.
Benefits
- Competitive base salary with equity in a unicorn-stage company.
- The company pays 100% of medical, dental, and vision premiums for employees and dependents.
- 401(k) matching of up to 4% of base pay.
- Unlimited paid time off and company-wide Refill Days throughout the year.
- Full-time, hybrid work arrangement in Boston.
Categories
About Liquid AI
We build efficient general-purpose AI at every scale.
