10 months ago
San Francisco, CA, USA or New York, NY, USAMid Level
H1B Sponsor
Base Salary
$180k - $360k/yr
Responsibilities
- Design, build, and operate the Model APIs surface, including structured outputs, tool and function calling, and multimodal serving
- Profile and optimize TensorRT-LLM kernels and CUDA execution, implement custom CUDA operators, tune memory allocation, and optimize multi-GPU communication
- Improve model-serving runtimes through speculative decoding, guided generation, quantization, batching, KV-cache reuse, and custom scheduling and routing
- Build benchmarking frameworks covering model architectures, batch sizes, sequence lengths, and hardware configurations
- Instrument metrics, traces, and logs to measure speed, reliability, and quality
- Implement API versioning, validation, usage metering, quotas, and authentication
- Collaborate with other teams to deliver robust, developer-friendly model-serving experiences
Requirements
- At least 3 years of experience building and operating distributed systems or large-scale APIs
- Proven experience owning low-latency, reliable backend services involving rate limiting, authentication, quotas, metering, and migrations
- Performance-focused infrastructure experience with profiling, tracing, capacity planning, and SLO management
- Ability to debug complex systems from runtime internals through GPU execution traces
- Strong written communication and ability to create clear design documents and collaborate across functions
- Experience with LLM runtimes such as vLLM, SGLang, or TensorRT-LLM, or contributions to open-source inference engines such as vLLM, TensorRT-LLM, SGLang, or TGI, is preferred
- Knowledge of Kubernetes, service meshes, API gateways, or distributed scheduling is preferred
- Background in developer-facing infrastructure or open-source APIs is preferred
- ML experience is a plus but not required
Benefits
- 100% medical, dental, and vision insurance coverage for employees and dependents
- Flexible PTO and a company-wide Winter Break from Christmas Eve through New Year's Day
- Paid parental leave
- Fertility and family-building stipend through Carrot
- Company-facilitated 401(k)
- Meaningful equity and exposure to a variety of ML startups
Tech Stack
About Baseten
Inference is everything. Baseten is an AI infrastructure platform giving you the tooling, expertise, and hardware needed to bring great AI products to market - fast. Our proprietary Inference Stack utilizes the cutting-edge of performance research combined with highly performant and reliable infrastructure to give you out-of-the-box global availability with 99.99% of uptime.
