3 months ago
Bellevue, WA, USA or Sunnyvale, CA, USASenior / Staff+
H1B sponsor
Base Salary
$182k - $242k/yr
Responsibilities
- Contribute to the Applied Training roadmap and determine which capabilities unlock important customer workloads.
- Work directly with customers and CoreWeave teams building compute, storage, networking, and other cloud-native primitives.
- Design and build research cluster experiences including CLI tools, job configuration schemas, Kubernetes operators, and daemons.
- Build infrastructure for code distribution, checkpoint-triggered evaluation, cross-cluster scheduling, and programmatic job control.
- Own the Python SDK and enable reinforcement-learning training runs to launch thousands of isolated containers for agent rollouts and benchmarks.
- Document how to run popular open-source training frameworks on CoreWeave.
- Study customer supercomputing stacks and apply those insights to the platform.
Requirements
- 8–12+ years building distributed systems, ML infrastructure, or developer platforms.
- Hands-on Kubernetes experience with custom controllers, operators, scheduling, CRDs, and workload orchestration at scale.
- Understanding of distributed training jobs, scheduling, rank initialization, and failure modes at scale.
- Experience shipping production infrastructure that people rely on daily.
- Strong communication skills and the ability to translate customer and researcher needs into system designs.
Benefits
- Medical, dental, and vision insurance fully paid by CoreWeave; company-paid life insurance; voluntary supplemental life insurance; short- and long-term disability insurance; Flexible Spending Account; Health Savings Account; tuition reimbursement; Employee Stock Purchase Program eligibility; Spring Health mental wellness benefits; Carrot family-forming support; paid parental leave; Kinside childcare support; 401(k) with employer match; flexible PTO; catered lunch at office and data center locations; casual work environment.
Tech Stack
Categories
About CoreWeave
CoreWeave provides a GPU-accelerated cloud for AI training and inference, VFX, and rendering, with bare-metal instances, Kubernetes orchestration, and managed services to scale workloads. It sells on-demand and reserved capacity to AI labs, startups, and enterprises, and offers SaaS tools and hands-on support for deployment. Founded in 2017 and headquartered in New York, it is publicly traded on Nasdaq under the ticker CRWV.
