
Inference Performance Engineer
Material Group4 months ago
Responsibilities
- Build and improve the inference runtime.
- Design scheduling, continuous batching, KV caching, and prefill/decode disaggregation.
- Implement low-precision kernels and speculative decoding.
- Optimize throughput, latency, and cost per token.
- Collaborate with hardware teams on kernels, operators, and graph optimizations.
- Own the OpenAI-compatible API surface and serving protocol.
- Build benchmarking, profiling, and regression infrastructure.
Requirements
- Bachelor's degree in computer science, electrical engineering, or a related field, or equivalent experience.
- Software engineering experience with Rust, Go, Python, or C++.
- Understanding of concurrency, memory, tail latency, transformers, attention, KV cache, batching, speculative decoding, and quantization.
- Experience with model-serving frameworks such as vLLM, TGI, SGLang, TensorRT-LLM, llama.cpp, or custom runtimes.
- GPU or ASIC programming experience with CUDA, ROCm, Triton, or vendor-native toolchains.
- Experience with low-precision inference using FP8, FP4, or INT4.
- Profiling and benchmarking experience with Nsight, perf, or custom harnesses.
Benefits
- Top-tier compensation structured to recognize and retain the best talent, plus meaningful equity.
- Comprehensive medical, dental, vision, life, and disability insurance.
- Parental leave for all new parents, including adoptive and surrogate journeys.
- Flexible PTO and paid holidays.
- Relocation support.