5 months ago
Base Salary
$188k - $275k/yr
Responsibilities
- Lead architecture and cross-team design initiatives spanning multiple services and teams
- Build and operate CoreWeave’s Kubernetes-native inference platform
- Optimize inference latency, throughput, GPU utilization, batching, scheduling, and memory usage
- Improve system reliability and performance across large-scale real-time inference systems
- Design and operate distributed, low-latency, high-throughput infrastructure
- Drive engineering direction through hands-on technical leadership and organization-wide influence
Requirements
- 8–12+ years of experience building and operating large-scale distributed systems or cloud platforms
- Experience leading cross-team technical initiatives affecting multiple services or organizations
- Strong programming skills in Go, Python, or C++
- Deep production-scale Kubernetes expertise, including orchestration, scheduling, and service design
- Strong understanding of distributed systems, networking, and performance optimization
- Experience designing and operating low-latency, high-throughput systems with strict P95/P99 latency requirements
- Hands-on experience with inference systems, batching or micro-batching, caching, and memory optimization
- Experience using metrics-driven approaches to improve latency, throughput, and utilization
- Familiarity with mixed precision, including BF16 and FP8, and streaming inference workloads
- Preferred experience with vLLM, Triton, TensorRT-LLM, Ray Serve, or TorchServe
- Preferred experience with GPU systems and optimization, including CUDA, NCCL, RDMA, NUMA, and GPU interconnects
- Preferred experience leading multi-team or organization-level technical initiatives
- Exposure to large-scale AI/ML infrastructure or hyperscale cloud environments
- Applicants must satisfy applicable U.S. export-control access or authorization requirements
Benefits
- Medical, dental, and vision insurance fully paid by CoreWeave
- Company-paid life insurance and voluntary supplemental life insurance
- Short- and long-term disability insurance
- Flexible Spending Account and Health Savings Account
- Tuition reimbursement
- Employee Stock Purchase Program eligibility
- Mental wellness benefits through Spring Health
- Family-forming support through Carrot
- Paid parental leave
- Flexible, full-service childcare support through Kinside
- 401(k) with employer match
- Flexible PTO
- Catered lunch at office and data center locations
- Casual work environment
- Hybrid work is prioritized; remote work may be considered for candidates more than 30 miles from an office
- Onboarding at a company hub within the first month and quarterly team gatherings
Tech Stack
Categories
About CoreWeave
CoreWeave provides a GPU-accelerated cloud for AI training and inference, VFX, and rendering, with bare-metal instances, Kubernetes orchestration, and managed services to scale workloads. It sells on-demand and reserved capacity to AI labs, startups, and enterprises, and offers SaaS tools and hands-on support for deployment. Founded in 2017 and headquartered in New York, it is publicly traded on Nasdaq under the ticker CRWV.
