5 months ago
Bellevue, WA, USA or Sunnyvale, CA, USAMid Level / Senior
H1B sponsor
Base Salary
$182k - $242k/yr
Responsibilities
- Develop and enhance Kubernetes-native benchmarking services measuring latency, throughput, jitter, and cost per request
- Implement and maintain end-to-end MLPerf Training and Inference benchmarking workflows, including workload setup, cluster configuration, and result validation
- Ingest, store, transform, and analyze performance events across global data centers
- Participate in design discussions and contribute to architecture decisions
- Break engineering work into milestones and deliver reliable, high-quality code
- Maintain reproducible and well-documented benchmarking processes
- Conduct code reviews and share engineering best practices
- Mentor junior engineers and review cross-team designs
- Improve coding and testing standards, latency, throughput, and reliability across multiple services
Requirements
- 3–5 years of experience building distributed systems, high-performance computing components, or cloud services
- Strong programming skills in Python or Go; C++ is a plus
- Understanding of networked systems and performance fundamentals
- Production experience with Kubernetes
- Familiarity with CI/CD and observability tools such as Prometheus, Grafana, and OpenTelemetry
- Exposure to performance-critical GPU systems or model-serving stacks such as llm-d, vLLM, TensorRT-LLM, and Megatron-LM
- Experience with time-series databases, LSM-based storage engines, or custom data pipelines is preferred
- Familiarity with MLPerf or other large-scale benchmarking frameworks is preferred
- Contributions to open-source projects such as llm-d, vLLM, or PyTorch are preferred
- Exposure to benchmarking GPU clusters or multi-region environments is preferred
- Background with CUDA kernels, NCCL/SHARP, RDMA/NUMA, or GPU interconnect topologies is preferred
- Must satisfy applicable export-control eligibility requirements
Benefits
- Medical, dental, and vision insurance fully paid by CoreWeave
- Company-paid life insurance, voluntary supplemental life insurance, and short- and long-term disability insurance
- Flexible Spending Account and Health Savings Account
- Tuition reimbursement and participation in the Employee Stock Purchase Program
- Mental wellness benefits through Spring Health and family-forming support through Carrot
- Paid parental leave and childcare support through Kinside
- 401(k) with employer match
- Flexible paid time off
- Catered lunch at office and data center locations
- Hybrid work environment, with remote work potentially available for candidates located more than 30 miles from an office
- Onboarding at a company hub within the first month and quarterly team gatherings
- Casual work environment and a culture focused on innovative disruption
About CoreWeave
CoreWeave provides a GPU-accelerated cloud for AI training and inference, VFX, and rendering, with bare-metal instances, Kubernetes orchestration, and managed services to scale workloads. It sells on-demand and reserved capacity to AI labs, startups, and enterprises, and offers SaaS tools and hands-on support for deployment. Founded in 2017 and headquartered in New York, it is publicly traded on Nasdaq under the ticker CRWV.
