5 months ago
Base Salary
$206k - $333k/yr
Responsibilities
- Define the multi-year benchmarking strategy and roadmap, prioritize models, workloads, and hardware tiers, and lead and mentor performance engineers and data analysts
- Establish governance for benchmark claims, including documented methodologies, versioning, reproducibility, and audit trails
- Lead end-to-end MLPerf Inference and Training submissions, including workload selection, cluster planning, runbooks, audits, and result publication
- Coordinate optimization efforts with NVIDIA across CUDA, cuDNN, TensorRT, TensorRT-LLM, Triton, and NCCL, and drive upstream fixes when needed
- Design a Kubernetes-native benchmarking service across SUNK, Kueue, and Kubeflow pipelines
- Measure latency, jitter, throughput, time-to-first-token, startup performance, and cost across models, precisions, batch sizes, and GPU types
- Maintain representative benchmarking scenarios and datasets and automate comparisons across software releases and hardware generations
- Build CI/CD pipelines and Kubernetes controllers or operators to schedule benchmarks at scale
- Integrate benchmarking systems with Prometheus, Grafana, OpenTelemetry, and results warehouses
- Implement supply-chain integrity for benchmark artifacts using SBOMs and Cosign signatures
- Partner with NVIDIA, ISVs, and open-source projects to co-develop optimizations and upstream improvements
- Provide authoritative performance data for RFPs and competitive evaluations and brief analysts and press
Requirements
- 10+ years building distributed systems or HPC/cloud services, with deep expertise in large-scale ML training or comparable high-performance workloads
- Proven experience architecting or building planet-scale data systems such as telemetry platforms, observability stacks, cloud data warehouses, or large-scale OLAP engines
- Deep understanding of GPU performance, including CUDA, NCCL, RDMA, NVLink/PCIe, and memory bandwidth
- Expertise with model-serving stacks such as Triton, vLLM, TensorRT-LLM, and TorchServe
- Experience with distributed training frameworks including PyTorch FSDP, DeepSpeed, and Megatron-LM
- Proficiency with Kubernetes and ML control planes, with production familiarity with SUNK, Kueue, and Kubeflow
- Excellent communication skills for working with executives, customers, auditors, and open-source communities
- Preferred experience with time-series databases, LSM trees, or custom storage engine development
- Preferred experience running MLPerf submissions or equivalent audited benchmarks at scale
- Preferred contributions to MLPerf, Triton, vLLM, PyTorch, KServe, or similar open-source projects
- Preferred experience benchmarking multi-region fleets and clusters containing thousands of GPUs
- Preferred publications or talks on ML performance, latency engineering, or large-scale benchmarking methodology
Benefits
- Medical, dental, and vision insurance fully paid by CoreWeave
- Company-paid life insurance and voluntary supplemental life insurance
- Short- and long-term disability insurance
- Flexible Spending Account and Health Savings Account
- Tuition reimbursement
- Employee Stock Purchase Program eligibility
- Mental wellness benefits through Spring Health
- Family-forming support through Carrot
- Paid parental leave
- Flexible childcare support through Kinside
- 401(k) with employer match
- Flexible paid time off
- Catered lunch at office and data center locations
- Hybrid work is prioritized; remote work may be considered for candidates more than 30 miles from an office
- Onboarding at a company hub within the first month and quarterly team gatherings
- Casual work environment focused on innovative disruption
Tech Stack
About CoreWeave
CoreWeave provides a GPU-accelerated cloud for AI training and inference, VFX, and rendering, with bare-metal instances, Kubernetes orchestration, and managed services to scale workloads. It sells on-demand and reserved capacity to AI labs, startups, and enterprises, and offers SaaS tools and hands-on support for deployment. Founded in 2017 and headquartered in New York, it is publicly traded on Nasdaq under the ticker CRWV.
