5 months ago
Base Salary
$165k - $242k/yr
Responsibilities
- Lead design reviews, drive architecture, and decompose multi-service initiatives into milestones
- Own an area spanning multiple services and teams, such as request routing, adaptive scheduling, cost-per-token analytics, or GPU resource isolation
- Define and own SLIs/SLOs and ensure post-incident actions improve reliability release over release
- Implement and quantify advanced inference optimizations including micro-batch schedulers, speculative decoding, and KV-cache reuse
- Plan capacity, improve autoscaling policies, and develop graceful degradation, rollback, and traffic-shift strategies
- Mentor IC1 and IC2 engineers, review cross-team designs, and raise coding and testing standards
- Partner with product, orchestration, and hardware teams to meet strict P99 service-level objectives
Requirements
- Approximately 5–8 years of industry experience building distributed systems or cloud services
- Strong coding skills in Python or Go; C++ is a plus
- Deep familiarity with networked systems and performance optimization
- Experience developing and tuning CUDA kernels, reducing model latency, and improving compute and memory bandwidth utilization
- Production-scale Kubernetes experience and familiarity with observability stacks including Prometheus, Grafana, and OpenTelemetry
- Knowledge of inference internals such as batching, caching, mixed precision with BF16/FP8, and streaming token delivery
- Track record of improving P95/P99 tail latency and service reliability through metrics-driven work
- Preferred contributions to inference frameworks such as vLLM, Triton, TensorRT-LLM, Ray Serve, or TorchServe
- Preferred experience with CUDA kernels, NCCL, SHARP, RDMA, NUMA, or GPU interconnect topologies
- Preferred experience leading multi-team initiatives or partnering with customers on mission-critical launches
Benefits
- Medical, dental, and vision insurance fully paid by CoreWeave
- Company-paid life insurance plus voluntary supplemental life insurance
- Short- and long-term disability insurance, Flexible Spending Account, and Health Savings Account
- Tuition reimbursement and eligibility to participate in the Employee Stock Purchase Program
- Mental wellness benefits through Spring Health and family-forming support through Carrot
- Paid parental leave and flexible full-service childcare support through Kinside
- 401(k) with employer match and flexible PTO
- Catered lunch at office and data center locations
- Hybrid work environment, with remote work potentially available for candidates located more than 30 miles from an office
- Onboarding at a company hub within the first month and quarterly team gatherings
Tech Stack
Categories
About CoreWeave
CoreWeave provides a GPU-accelerated cloud for AI training and inference, VFX, and rendering, with bare-metal instances, Kubernetes orchestration, and managed services to scale workloads. It sells on-demand and reserved capacity to AI labs, startups, and enterprises, and offers SaaS tools and hands-on support for deployment. Founded in 2017 and headquartered in New York, it is publicly traded on Nasdaq under the ticker CRWV.
