25 days ago
Remote, United StatesSenior
Responsibilities
- Stand up and optimize GPU infrastructure, including B300 nodes in owned data centers.
- Improve TTFT, TPOT, throughput, and cost per token for LLM inference workloads.
- Build reproducible benchmarking harnesses across inference engines, quantization schemes, parallelism strategies, workloads, and GPU SKUs.
- Optimize multivariate inference load-balancing algorithms within the inference routing system.
- Evaluate custom CUDA and Triton kernels, attention variants, quantization schemes, and compilation improvements.
- Evaluate emerging inference hardware, including FPGAs, ASICs, and custom silicon.
Requirements
- 5+ years of experience in performance optimization or HPC, with deep GPU architecture and parallel programming knowledge.
- Hands-on production experience with at least one high-volume LLM inference engine, such as vLLM or SGLang.
- Experience with continuous batching, PagedAttention, KV cache management, speculative decoding, quantization, CUDA graphs, and torch.compile.
- Experience with tensor, pipeline, and MoE parallelism in multi-GPU and multi-node environments.
- Fluency with Nsight Systems, Nsight Compute, and PyTorch Profiler for GPU profiling.
- Proficiency in Python, Rust, or Go; C++ and CUDA are strong pluses.
- Bonus qualifications include custom Triton kernels, diffusion or image model inference optimization, and contributions to open-source inference frameworks.
- Passion for AI and/or crypto is encouraged, though applicants do not need to match every qualification.
Benefits
- Remote position in the USA, with exceptional candidates outside the USA considered.
- This is an application through Dragonfly’s talent network for a portfolio company rather than an internal Dragonfly role.
- Dragonfly may facilitate a warm introduction to the portfolio company and retain candidates for future opportunities.
