Baseten

Software Engineer - GPU Kernels

Baseten
Apply
1 year ago
San Francisco, CA, USA or New York, NY, USAMid Level
H1B sponsor

Base Salary

$180k - $360k/yr

Responsibilities

  • Design and implement high-performance GPU kernels for matrix multiplications, attention mechanisms, mixture-of-experts routing, and other key machine-learning operations.
  • Optimize code using CUDA, PTX assembly, and architecture-specific techniques.
  • Apply memory coalescing, warp-level programming, tensor core acceleration, and compute/memory overlap.
  • Implement quantization, sparsity, and compute/communication overlap features.
  • Identify and resolve performance bottlenecks using Nsight Systems, Nsight Compute, and Torch Profiler.
  • Collaborate with research teams to productionize theoretical advancements.
  • Contribute to internal and open-source GPU libraries.
  • Present technical contributions at industry conferences such as NVIDIA GTC and AWS re:Invent.

Requirements

  • Strong understanding of GPU architecture and programming paradigms, including memory hierarchy, thread/block/grid organization, synchronization, and race-condition mitigation.
  • Proficiency in C++ and GPU performance profiling tools.
  • Knowledge of the CUDA C++ API, memory access patterns, bandwidth optimization, numerical precision, quantization strategies, tensor cores, and asynchronous operations.
  • Experience with Transformer models and attention optimization, such as Flash Attention, is preferred.
  • Familiarity with Cutlass, Triton, Thrust, or CUB is preferred.
  • Background in GEMM tuning and distributed or multi-GPU compute is preferred.
  • Open-source GPU contributions and research publications or conference presentations on GPU performance are preferred.

Benefits

  • 100% medical, dental, and vision insurance coverage for employees and dependents.
  • Flexible paid time off and a company-wide Winter Break from Christmas Eve through New Year's Day.
  • Paid parental leave.
  • Fertility and family-building stipend through Carrot.
  • Company-facilitated 401(k).
  • Meaningful equity and exposure to a variety of ML startups.

Tech Stack

Categories

Baseten

About Baseten

201-500 employees

Baseten builds an AI inference platform that provides tooling, infrastructure, and hardware to deploy, scale, and serve machine-learning models in production. The company sells managed model serving and developer tooling to software teams at AI product companies, with customers including Notion, Abridge, Writer, and Cursor. Privately held and headquartered in San Francisco, it focuses on high-availability, globally distributed inference for production workloads.

Contact me