3 months ago
Base Salary
$188k - $275k/yr
Responsibilities
- Lead architecture and cross-team design initiatives spanning multiple services and teams
- Build and operate CoreWeave’s Kubernetes-native inference platform
- Optimize inference latency, throughput, GPU utilization, batching, scheduling, and memory usage
- Improve system reliability and performance across large-scale real-time inference systems
- Design and operate distributed, low-latency, high-throughput infrastructure
- Drive engineering direction through hands-on technical leadership and organization-wide influence
Requirements
- 8–12+ years of experience building and operating large-scale distributed systems or cloud platforms
- Experience leading cross-team technical initiatives affecting multiple services or organizations
- Strong programming skills in Go, Python, or C++
- Deep production-scale Kubernetes expertise, including orchestration, scheduling, and service design
- Strong understanding of distributed systems, networking, and performance optimization
- Experience designing and operating low-latency, high-throughput systems with strict P95/P99 latency requirements
- Hands-on experience with inference systems, batching or micro-batching, caching, and memory optimization
- Experience using metrics-driven approaches to improve latency, throughput, and utilization
- Familiarity with mixed precision, including BF16 and FP8, and streaming inference workloads
- Preferred experience with vLLM, Triton, TensorRT-LLM, Ray Serve, or TorchServe
- Preferred experience with GPU systems and optimization, including CUDA, NCCL, RDMA, NUMA, and GPU interconnects
- Preferred experience leading multi-team or organization-level technical initiatives
- Exposure to large-scale AI/ML infrastructure or hyperscale cloud environments
- Applicants must satisfy applicable U.S. export-control access or authorization requirements
Benefits
- Medical, dental, and vision insurance fully paid by CoreWeave
- Company-paid life insurance and voluntary supplemental life insurance
- Short- and long-term disability insurance
- Flexible Spending Account and Health Savings Account
- Tuition reimbursement
- Employee Stock Purchase Program eligibility
- Mental wellness benefits through Spring Health
- Family-forming support through Carrot
- Paid parental leave
- Flexible, full-service childcare support through Kinside
- 401(k) with employer match
- Flexible PTO
- Catered lunch at office and data center locations
- Casual work environment
- Hybrid work is prioritized; remote work may be considered for candidates more than 30 miles from an office
- Onboarding at a company hub within the first month and quarterly team gatherings
Tech Stack
Categories
About CoreWeave
CoreWeave is the Essential Cloud for AI. CoreWeave is a cloud purpose-built for scaling, supporting, and accelerating GenAI. We’re a comprehensive platform and strategic partner designed to tackle today—and tomorrow’s—challenges of deploying AI at scale. We manage the complexities of AI growth to make supercomputing accessible and push the limits of what’s possible. Our teams create modern solutions to support modern technology. Get the premier choice for working with GenAI workloads.
