1 year ago
San Francisco, CA, USA or New York, NY, USAMid Level
H1B Sponsor
Base Salary
$180k - $360k/yr
Responsibilities
- Design and implement high-performance GPU kernels for matrix multiplications, attention mechanisms, mixture-of-experts routing, and other key machine-learning operations.
- Optimize code using CUDA, PTX assembly, and architecture-specific techniques.
- Apply memory coalescing, warp-level programming, tensor core acceleration, and compute/memory overlap.
- Implement quantization, sparsity, and compute/communication overlap features.
- Identify and resolve performance bottlenecks using Nsight Systems, Nsight Compute, and Torch Profiler.
- Collaborate with research teams to productionize theoretical advancements.
- Contribute to internal and open-source GPU libraries.
- Present technical contributions at industry conferences such as NVIDIA GTC and AWS re:Invent.
Requirements
- Strong understanding of GPU architecture and programming paradigms, including memory hierarchy, thread/block/grid organization, synchronization, and race-condition mitigation.
- Proficiency in C++ and GPU performance profiling tools.
- Knowledge of the CUDA C++ API, memory access patterns, bandwidth optimization, numerical precision, quantization strategies, tensor cores, and asynchronous operations.
- Experience with Transformer models and attention optimization, such as Flash Attention, is preferred.
- Familiarity with Cutlass, Triton, Thrust, or CUB is preferred.
- Background in GEMM tuning and distributed or multi-GPU compute is preferred.
- Open-source GPU contributions and research publications or conference presentations on GPU performance are preferred.
Benefits
- 100% medical, dental, and vision insurance coverage for employees and dependents.
- Flexible paid time off and a company-wide Winter Break from Christmas Eve through New Year's Day.
- Paid parental leave.
- Fertility and family-building stipend through Carrot.
- Company-facilitated 401(k).
- Meaningful equity and exposure to a variety of ML startups.
About Baseten
Inference is everything. Baseten is an AI infrastructure platform giving you the tooling, expertise, and hardware needed to bring great AI products to market - fast. Our proprietary Inference Stack utilizes the cutting-edge of performance research combined with highly performant and reliable infrastructure to give you out-of-the-box global availability with 99.99% of uptime.
