3 months ago
Bellevue, WA, USA or Sunnyvale, CA, USAMid Level / Senior
H1B Sponsor
Base Salary
$182k - $242k/yr
Responsibilities
- Develop and enhance Kubernetes-native benchmarking services measuring latency, throughput, jitter, and cost per request
- Implement and maintain end-to-end MLPerf Training and Inference benchmarking workflows, including workload setup, cluster configuration, and result validation
- Ingest, store, transform, and analyze performance events across global data centers
- Participate in design discussions and contribute to architecture decisions
- Break engineering work into milestones and deliver reliable, high-quality code
- Maintain reproducible and well-documented benchmarking processes
- Conduct code reviews and share engineering best practices
- Mentor junior engineers and review cross-team designs
- Improve coding and testing standards, latency, throughput, and reliability across multiple services
Requirements
- 3–5 years of experience building distributed systems, high-performance computing components, or cloud services
- Strong programming skills in Python or Go; C++ is a plus
- Understanding of networked systems and performance fundamentals
- Production experience with Kubernetes
- Familiarity with CI/CD and observability tools such as Prometheus, Grafana, and OpenTelemetry
- Exposure to performance-critical GPU systems or model-serving stacks such as llm-d, vLLM, TensorRT-LLM, and Megatron-LM
- Experience with time-series databases, LSM-based storage engines, or custom data pipelines is preferred
- Familiarity with MLPerf or other large-scale benchmarking frameworks is preferred
- Contributions to open-source projects such as llm-d, vLLM, or PyTorch are preferred
- Exposure to benchmarking GPU clusters or multi-region environments is preferred
- Background with CUDA kernels, NCCL/SHARP, RDMA/NUMA, or GPU interconnect topologies is preferred
- Must satisfy applicable export-control eligibility requirements
Benefits
- Medical, dental, and vision insurance fully paid by CoreWeave
- Company-paid life insurance, voluntary supplemental life insurance, and short- and long-term disability insurance
- Flexible Spending Account and Health Savings Account
- Tuition reimbursement and participation in the Employee Stock Purchase Program
- Mental wellness benefits through Spring Health and family-forming support through Carrot
- Paid parental leave and childcare support through Kinside
- 401(k) with employer match
- Flexible paid time off
- Catered lunch at office and data center locations
- Hybrid work environment, with remote work potentially available for candidates located more than 30 miles from an office
- Onboarding at a company hub within the first month and quarterly team gatherings
- Casual work environment and a culture focused on innovative disruption
About CoreWeave
CoreWeave is the Essential Cloud for AI. CoreWeave is a cloud purpose-built for scaling, supporting, and accelerating GenAI. We’re a comprehensive platform and strategic partner designed to tackle today—and tomorrow’s—challenges of deploying AI at scale. We manage the complexities of AI growth to make supercomputing accessible and push the limits of what’s possible. Our teams create modern solutions to support modern technology. Get the premier choice for working with GenAI workloads.
