Baseten

Software Engineer - GPU Kernels

Baseten
Apply
1 year ago
San Francisco, CA, USA or New York, NY, USAMid Level
H1B Sponsor

Base Salary

$180k - $360k/yr

Responsibilities

  • Design and implement high-performance GPU kernels for matrix multiplications, attention mechanisms, mixture-of-experts routing, and other key machine-learning operations.
  • Optimize code using CUDA, PTX assembly, and architecture-specific techniques.
  • Apply memory coalescing, warp-level programming, tensor core acceleration, and compute/memory overlap.
  • Implement quantization, sparsity, and compute/communication overlap features.
  • Identify and resolve performance bottlenecks using Nsight Systems, Nsight Compute, and Torch Profiler.
  • Collaborate with research teams to productionize theoretical advancements.
  • Contribute to internal and open-source GPU libraries.
  • Present technical contributions at industry conferences such as NVIDIA GTC and AWS re:Invent.

Requirements

  • Strong understanding of GPU architecture and programming paradigms, including memory hierarchy, thread/block/grid organization, synchronization, and race-condition mitigation.
  • Proficiency in C++ and GPU performance profiling tools.
  • Knowledge of the CUDA C++ API, memory access patterns, bandwidth optimization, numerical precision, quantization strategies, tensor cores, and asynchronous operations.
  • Experience with Transformer models and attention optimization, such as Flash Attention, is preferred.
  • Familiarity with Cutlass, Triton, Thrust, or CUB is preferred.
  • Background in GEMM tuning and distributed or multi-GPU compute is preferred.
  • Open-source GPU contributions and research publications or conference presentations on GPU performance are preferred.

Benefits

  • 100% medical, dental, and vision insurance coverage for employees and dependents.
  • Flexible paid time off and a company-wide Winter Break from Christmas Eve through New Year's Day.
  • Paid parental leave.
  • Fertility and family-building stipend through Carrot.
  • Company-facilitated 401(k).
  • Meaningful equity and exposure to a variety of ML startups.

Tech Stack

Categories

Baseten

About Baseten

201-500 employees

Inference is everything. Baseten is an AI infrastructure platform giving you the tooling, expertise, and hardware needed to bring great AI products to market - fast. Our proprietary Inference Stack utilizes the cutting-edge of performance research combined with highly performant and reliable infrastructure to give you out-of-the-box global availability with 99.99% of uptime.