over 2 years ago
Base Salary
$180k - $360k/yr
Responsibilities
- Implement, refine, and productionize ML inference techniques including quantization, speculative decoding, KV cache reuse, chunked prefill, and LoRA
- Debug ML performance issues by analyzing TensorRT, PyTorch, TensorRT-LLM, vLLM, SGLang, CUDA, and other codebases
- Apply and scale optimization techniques across a wide range of ML models, especially large language models
- Collaborate with a diverse team to design and implement innovative solutions
- Own projects from idea through production
Requirements
- Bachelor's, Master's, or Ph.D. degree in Computer Science, Engineering, Mathematics, or a related field
- Experience with a general-purpose programming language such as Python or C++
- Familiarity with LLM optimization techniques including quantization, speculative decoding, and continuous batching
- Strong familiarity with ML libraries, especially PyTorch, TensorRT, or TensorRT-LLM
- Demonstrated interest and experience in LLMs
- Deep understanding of GPU architecture
- Proficiency enhancing software-system performance, particularly for large language models, is preferred
- Experience with CUDA or similar technologies is preferred
- Deep understanding of software engineering principles and a proven track record developing and deploying AI/ML inference solutions is preferred
- Experience with Docker and Kubernetes is preferred
Benefits
- Competitive compensation including meaningful equity
- 100% coverage of medical, dental, and vision insurance for employees and dependents
- Flexible paid time off, including a company-wide Winter Break from Christmas Eve through New Year's Day
- Paid parental leave
- Fertility and family-building stipend through Carrot
- Company-facilitated 401(k)
- Exposure to a variety of ML startups and learning and networking opportunities
Tech Stack
Categories
About Baseten
Inference is everything. Baseten is an AI infrastructure platform giving you the tooling, expertise, and hardware needed to bring great AI products to market - fast. Our proprietary Inference Stack utilizes the cutting-edge of performance research combined with highly performant and reliable infrastructure to give you out-of-the-box global availability with 99.99% of uptime.
