
AI Inference Engineer
Quadric, Inc10 months ago
Burlingame, CA, USASenior
Base Salary
$110k - $270k/yr
Responsibilities
- Quantize, prune, and convert AI models for deployment.
- Port models to the Quadric platform using the Quadric toolchain.
- Optimize inference deployments for latency and speed.
- Benchmark and profile model performance and accuracy.
- Collaborate across related areas of the AI inference stack.
- Develop tools to scale and accelerate model deployment.
- Improve the SDK and runtime.
- Provide technical support and documentation to customers and the developer community.
Requirements
- Bachelor’s or master’s degree in Computer Science and/or Electrical Engineering.
- 5+ years of experience with AI/LLM model inference and deployment frameworks or tools.
- Experience with model quantization, including PTQ and QAT.
- Experience with model accuracy measures and inference performance profiling.
- Experience with at least one of ONNX Runtime, PyTorch, vLLM, Hugging Face Transformers, Neural Compressor, or llama.cpp.
- Proficiency in C/C++ and Python.
- Demonstrated problem-solving, debugging, and communication capability.
Benefits
- Competitive base salary, meaningful equity, and a discretionary annual performance bonus as applicable.
- Medical, dental, and vision plans beginning on day one.
- 401(k) retirement plan.
- Flexible unlimited paid time off.
- Company-provided lunches and a stocked kitchen when working in the office.
- Hybrid schedule with at least two in-office days per week at the Burlingame office and occasional additional onsite days as needed.
- Commuting support, including monthly parking or Caltrain passes.
- Downtown Burlingame office within walking distance of Caltrain and near shops, cafes, and local amenities.
- Collaborative, politics-free work environment and opportunities to build long-term career relationships.