over 2 years ago
Base Salary
$180k - $360k/yr
Responsibilities
- Implement, refine, and productionize ML inference techniques including quantization, speculative decoding, KV cache reuse, chunked prefill, and LoRA
- Debug ML performance issues by analyzing TensorRT, PyTorch, TensorRT-LLM, vLLM, SGLang, CUDA, and other codebases
- Apply and scale optimization techniques across a wide range of ML models, especially large language models
- Collaborate with a diverse team to design and implement innovative solutions
- Own projects from idea through production
Requirements
- Bachelor's, Master's, or Ph.D. degree in Computer Science, Engineering, Mathematics, or a related field
- Experience with a general-purpose programming language such as Python or C++
- Familiarity with LLM optimization techniques including quantization, speculative decoding, and continuous batching
- Strong familiarity with ML libraries, especially PyTorch, TensorRT, or TensorRT-LLM
- Demonstrated interest and experience in LLMs
- Deep understanding of GPU architecture
- Proficiency enhancing software-system performance, particularly for large language models, is preferred
- Experience with CUDA or similar technologies is preferred
- Deep understanding of software engineering principles and a proven track record developing and deploying AI/ML inference solutions is preferred
- Experience with Docker and Kubernetes is preferred
Benefits
- Competitive compensation including meaningful equity
- 100% coverage of medical, dental, and vision insurance for employees and dependents
- Flexible paid time off, including a company-wide Winter Break from Christmas Eve through New Year's Day
- Paid parental leave
- Fertility and family-building stipend through Carrot
- Company-facilitated 401(k)
- Exposure to a variety of ML startups and learning and networking opportunities
Tech Stack
Categories
About Baseten
Baseten builds an AI inference platform that provides tooling, infrastructure, and hardware to deploy, scale, and serve machine-learning models in production. The company sells managed model serving and developer tooling to software teams at AI product companies, with customers including Notion, Abridge, Writer, and Cursor. Privately held and headquartered in San Francisco, it focuses on high-availability, globally distributed inference for production workloads.
