CoreWeave

Principal Engineer - Perf and Benchmarking

CoreWeave
Apply
5 months ago
Bellevue, WA, USA or Sunnyvale, CA, USAStaff+
H1B sponsor

Base Salary

$206k - $333k/yr

Responsibilities

  • Define the multi-year benchmarking strategy and roadmap, prioritize models, workloads, and hardware tiers, and lead and mentor performance engineers and data analysts
  • Establish governance for benchmark claims, including documented methodologies, versioning, reproducibility, and audit trails
  • Lead end-to-end MLPerf Inference and Training submissions, including workload selection, cluster planning, runbooks, audits, and result publication
  • Coordinate optimization efforts with NVIDIA across CUDA, cuDNN, TensorRT, TensorRT-LLM, Triton, and NCCL, and drive upstream fixes when needed
  • Design a Kubernetes-native benchmarking service across SUNK, Kueue, and Kubeflow pipelines
  • Measure latency, jitter, throughput, time-to-first-token, startup performance, and cost across models, precisions, batch sizes, and GPU types
  • Maintain representative benchmarking scenarios and datasets and automate comparisons across software releases and hardware generations
  • Build CI/CD pipelines and Kubernetes controllers or operators to schedule benchmarks at scale
  • Integrate benchmarking systems with Prometheus, Grafana, OpenTelemetry, and results warehouses
  • Implement supply-chain integrity for benchmark artifacts using SBOMs and Cosign signatures
  • Partner with NVIDIA, ISVs, and open-source projects to co-develop optimizations and upstream improvements
  • Provide authoritative performance data for RFPs and competitive evaluations and brief analysts and press

Requirements

  • 10+ years building distributed systems or HPC/cloud services, with deep expertise in large-scale ML training or comparable high-performance workloads
  • Proven experience architecting or building planet-scale data systems such as telemetry platforms, observability stacks, cloud data warehouses, or large-scale OLAP engines
  • Deep understanding of GPU performance, including CUDA, NCCL, RDMA, NVLink/PCIe, and memory bandwidth
  • Expertise with model-serving stacks such as Triton, vLLM, TensorRT-LLM, and TorchServe
  • Experience with distributed training frameworks including PyTorch FSDP, DeepSpeed, and Megatron-LM
  • Proficiency with Kubernetes and ML control planes, with production familiarity with SUNK, Kueue, and Kubeflow
  • Excellent communication skills for working with executives, customers, auditors, and open-source communities
  • Preferred experience with time-series databases, LSM trees, or custom storage engine development
  • Preferred experience running MLPerf submissions or equivalent audited benchmarks at scale
  • Preferred contributions to MLPerf, Triton, vLLM, PyTorch, KServe, or similar open-source projects
  • Preferred experience benchmarking multi-region fleets and clusters containing thousands of GPUs
  • Preferred publications or talks on ML performance, latency engineering, or large-scale benchmarking methodology

Benefits

  • Medical, dental, and vision insurance fully paid by CoreWeave
  • Company-paid life insurance and voluntary supplemental life insurance
  • Short- and long-term disability insurance
  • Flexible Spending Account and Health Savings Account
  • Tuition reimbursement
  • Employee Stock Purchase Program eligibility
  • Mental wellness benefits through Spring Health
  • Family-forming support through Carrot
  • Paid parental leave
  • Flexible childcare support through Kinside
  • 401(k) with employer match
  • Flexible paid time off
  • Catered lunch at office and data center locations
  • Hybrid work is prioritized; remote work may be considered for candidates more than 30 miles from an office
  • Onboarding at a company hub within the first month and quarterly team gatherings
  • Casual work environment focused on innovative disruption

Tech Stack

GrafanaKubernetesPrometheusPyTorch

Categories

BackendData EngineeringDevOps
CoreWeave

About CoreWeave

1,001-5,000 employees

CoreWeave provides a GPU-accelerated cloud for AI training and inference, VFX, and rendering, with bare-metal instances, Kubernetes orchestration, and managed services to scale workloads. It sells on-demand and reserved capacity to AI labs, startups, and enterprises, and offers SaaS tools and hands-on support for deployment. Founded in 2017 and headquartered in New York, it is publicly traded on Nasdaq under the ticker CRWV.

Contact me