1 month ago
Base Salary
$182k - $242k/yr
Responsibilities
- Author, profile, and optimize CUDA kernels for GEMMs, attention, MoE routing, quantization, KV-cache operations, and fused epilogues in LLM inference
- Tune GPU execution using tensor cores, occupancy, memory coalescing, shared-memory and register usage, and compute/data-movement overlap
- Use kernel-authoring DSLs and compilers to prototype and ship high-performance kernels
- Build reproducible microbenchmarks and roofline analyses and validate end-to-end latency and throughput gains across model-serving stacks
- Implement and maintain MLPerf Inference and Training benchmarking workflows, including workload setup, cluster configuration, runbooks, and result validation
- Lead design reviews, drive team architecture, and decompose multi-service work into milestones
- Mentor junior engineers, review cross-team designs, and improve coding and testing standards
- Maintain reproducible and well-documented benchmarking and kernel-optimization processes
Requirements
- At least 5 years of experience building high-performance computing, GPU/accelerator software, or performance-critical systems
- Required hands-on CUDA experience, including writing and optimizing custom kernels and fluency with the CUDA programming and memory model
- Deep understanding of GPU architecture and performance, including tensor cores, warp and occupancy tuning, memory hierarchy and bandwidth, NVLink, PCIe, and Nsight profiling
- Strong C++ and Python programming skills with comfort writing and reading low-level, performance-sensitive code
- Familiarity with vLLM, TensorRT-LLM, llm-d, SGLang, and inference-dominant kernels
- Strong communication and cross-functional collaboration skills
- Preferred experience with Triton or Mojo for custom GPU kernel authoring
- Preferred experience with CuTe DSL, JAX, Pallas, HIP, ROCm, NCCL, Google TPUs, Meta MTIA, KNYFE, Block DSL, Kubernetes, SUNK, Slurm, MLPerf, vLLM, SGLang, PyTorch, Triton, or CUTLASS
- Experience with MLPerf submissions or similar large-scale audited benchmarks is preferred
- Contributions to open-source projects such as vLLM, SGLang, PyTorch, Triton, or CUTLASS are preferred
- Must meet applicable US export-control eligibility requirements for access to controlled information
Benefits
- Medical, dental, and vision insurance fully paid by CoreWeave for US-based employees
- Company-paid life insurance and voluntary supplemental life insurance
- Short- and long-term disability insurance
- Flexible Spending Account and Health Savings Account
- Tuition reimbursement
- Employee Stock Purchase Program participation
- Mental wellness benefits through Spring Health
- Family-forming support through Carrot
- Paid parental leave
- Flexible, full-service childcare support through Kinside
- 401(k) with an employer match
- Flexible PTO
- Catered lunch at office and data center locations
- Casual work environment
- Export-controlled information access requirements apply, subject to US export regulations
Tech Stack
Categories
About CoreWeave
CoreWeave is the Essential Cloud for AI. CoreWeave is a cloud purpose-built for scaling, supporting, and accelerating GenAI. We’re a comprehensive platform and strategic partner designed to tackle today—and tomorrow’s—challenges of deploying AI at scale. We manage the complexities of AI growth to make supercomputing accessible and push the limits of what’s possible. Our teams create modern solutions to support modern technology. Get the premier choice for working with GenAI workloads.
