Baseten

Software Engineer — GPU Networking & Distributed Systems

Baseten
Apply
7 months ago
San Francisco, CA, USA or New York, NY, USASenior
H1B sponsor

Base Salary

$165k - $330k/yr

Responsibilities

  • Integrate RDMA, RoCE, and InfiniBand capabilities directly into the inference stack.
  • Implement and tune networking layers for Disaggregated KV Cache Offload and Wide Expert Parallelism across NVLink and InfiniBand.
  • Develop checkpointing and storage mechanisms to enable sub-10-second startup for trillion-parameter models.
  • Characterize and validate networking performance on H100, H200, B200, B300, and GB200/300 NVL72 clusters.
  • Write acceptance tests that validate hardware throughput and latency.
  • Build observability tools for visualizing packet flow, congestion, and effective bandwidth across GPU interconnects.
  • Optimize NCCL and NVSHMEM communication and potentially develop custom communication kernels to overlap computation and data transfer.

Requirements

  • Deep experience with high-performance networking protocols, including InfiniBand and RoCE v2.
  • Fluency in C++ or Python and the ability to bridge high-level software logic with hardware.
  • Deep understanding of memory hierarchies in modern NVIDIA architectures, including H100 and Blackwell.
  • Ability to work with TensorRT-LLM source code, write custom C++ or Python bindings, and debug NVLink topology issues.
  • Ability to evaluate when to use off-the-shelf solutions and when to build custom networking infrastructure.
  • Highly preferred: deep knowledge of NCCL, NVSHMEM, and UCX.
  • Highly preferred: experience with GPUDirect Storage or high-performance filesystems such as Weka or 3FS.
  • Highly preferred: familiarity with TensorRT-LLM, vLLM, or Sglang.
  • Highly preferred: experience running low-level benchmarks to qualify new hardware clusters.

Benefits

  • 100% medical, dental, and vision insurance coverage for employees and dependents.
  • Flexible PTO and a company-wide Winter Break, with offices closed from Christmas Eve through New Year's Day.
  • Paid parental leave.
  • Fertility and family-building stipend through Carrot.
  • Company-facilitated 401(k).
  • Exposure to a variety of machine learning startups and networking opportunities.
Baseten

About Baseten

201-500 employees

Baseten builds an AI inference platform that provides tooling, infrastructure, and hardware to deploy, scale, and serve machine-learning models in production. The company sells managed model serving and developer tooling to software teams at AI product companies, with customers including Notion, Abridge, Writer, and Cursor. Privately held and headquartered in San Francisco, it focuses on high-availability, globally distributed inference for production workloads.

Contact me