4 months ago
Base Salary
$165k - $242k/yr
Responsibilities
- Lead design reviews, drive architecture, and decompose multi-service initiatives into milestones
- Own an area spanning multiple services and teams, such as request routing, adaptive scheduling, cost-per-token analytics, or GPU resource isolation
- Define and own SLIs/SLOs and ensure post-incident actions improve reliability release over release
- Implement and quantify advanced inference optimizations including micro-batch schedulers, speculative decoding, and KV-cache reuse
- Plan capacity, improve autoscaling policies, and develop graceful degradation, rollback, and traffic-shift strategies
- Mentor IC1 and IC2 engineers, review cross-team designs, and raise coding and testing standards
- Partner with product, orchestration, and hardware teams to meet strict P99 service-level objectives
Requirements
- Approximately 5–8 years of industry experience building distributed systems or cloud services
- Strong coding skills in Python or Go; C++ is a plus
- Deep familiarity with networked systems and performance optimization
- Experience developing and tuning CUDA kernels, reducing model latency, and improving compute and memory bandwidth utilization
- Production-scale Kubernetes experience and familiarity with observability stacks including Prometheus, Grafana, and OpenTelemetry
- Knowledge of inference internals such as batching, caching, mixed precision with BF16/FP8, and streaming token delivery
- Track record of improving P95/P99 tail latency and service reliability through metrics-driven work
- Preferred contributions to inference frameworks such as vLLM, Triton, TensorRT-LLM, Ray Serve, or TorchServe
- Preferred experience with CUDA kernels, NCCL, SHARP, RDMA, NUMA, or GPU interconnect topologies
- Preferred experience leading multi-team initiatives or partnering with customers on mission-critical launches
Benefits
- Medical, dental, and vision insurance fully paid by CoreWeave
- Company-paid life insurance plus voluntary supplemental life insurance
- Short- and long-term disability insurance, Flexible Spending Account, and Health Savings Account
- Tuition reimbursement and eligibility to participate in the Employee Stock Purchase Program
- Mental wellness benefits through Spring Health and family-forming support through Carrot
- Paid parental leave and flexible full-service childcare support through Kinside
- 401(k) with employer match and flexible PTO
- Catered lunch at office and data center locations
- Hybrid work environment, with remote work potentially available for candidates located more than 30 miles from an office
- Onboarding at a company hub within the first month and quarterly team gatherings
Tech Stack
Categories
About CoreWeave
CoreWeave is the Essential Cloud for AI. CoreWeave is a cloud purpose-built for scaling, supporting, and accelerating GenAI. We’re a comprehensive platform and strategic partner designed to tackle today—and tomorrow’s—challenges of deploying AI at scale. We manage the complexities of AI growth to make supercomputing accessible and push the limits of what’s possible. Our teams create modern solutions to support modern technology. Get the premier choice for working with GenAI workloads.
