6 months ago
Base Salary
$165k - $330k/yr
Responsibilities
- Integrate RDMA, RoCE, and InfiniBand capabilities directly into the inference stack.
- Implement and tune networking layers for Disaggregated KV Cache Offload and Wide Expert Parallelism across NVLink and InfiniBand.
- Develop checkpointing and storage mechanisms to enable sub-10-second startup for trillion-parameter models.
- Characterize and validate networking performance on H100, H200, B200, B300, and GB200/300 NVL72 clusters.
- Write acceptance tests that validate hardware throughput and latency.
- Build observability tools for visualizing packet flow, congestion, and effective bandwidth across GPU interconnects.
- Optimize NCCL and NVSHMEM communication and potentially develop custom communication kernels to overlap computation and data transfer.
Requirements
- Deep experience with high-performance networking protocols, including InfiniBand and RoCE v2.
- Fluency in C++ or Python and the ability to bridge high-level software logic with hardware.
- Deep understanding of memory hierarchies in modern NVIDIA architectures, including H100 and Blackwell.
- Ability to work with TensorRT-LLM source code, write custom C++ or Python bindings, and debug NVLink topology issues.
- Ability to evaluate when to use off-the-shelf solutions and when to build custom networking infrastructure.
- Highly preferred: deep knowledge of NCCL, NVSHMEM, and UCX.
- Highly preferred: experience with GPUDirect Storage or high-performance filesystems such as Weka or 3FS.
- Highly preferred: familiarity with TensorRT-LLM, vLLM, or Sglang.
- Highly preferred: experience running low-level benchmarks to qualify new hardware clusters.
Benefits
- 100% medical, dental, and vision insurance coverage for employees and dependents.
- Flexible PTO and a company-wide Winter Break, with offices closed from Christmas Eve through New Year's Day.
- Paid parental leave.
- Fertility and family-building stipend through Carrot.
- Company-facilitated 401(k).
- Exposure to a variety of machine learning startups and networking opportunities.
Tech Stack
About Baseten
Inference is everything. Baseten is an AI infrastructure platform giving you the tooling, expertise, and hardware needed to bring great AI products to market - fast. Our proprietary Inference Stack utilizes the cutting-edge of performance research combined with highly performant and reliable infrastructure to give you out-of-the-box global availability with 99.99% of uptime.
