5 months ago
Base Salary
$182k - $242k/yr
Responsibilities
- Design, build, and operate Go-based services managing large-scale GPU data center infrastructure
- Automate data center bring-up, hardware discovery, health monitoring, remediation, and production operations
- Develop APIs, services, and workflows for BMCs, firmware state, server health, and rack-level infrastructure
- Improve observability, alerting, and operational tooling for rapid production issue detection and resolution
- Convert incidents and hardware failure modes into software improvements that increase platform resilience
- Collaborate with hardware-adjacent, infrastructure, operations, and software teams on fleet-scale systems
Requirements
- At least 5 years of experience building and operating infrastructure or backend systems
- Bachelor’s or Master’s degree in Computer Science or a related field, or equivalent practical experience
- Strong proficiency in Go for production services and tools
- Experience designing and building gRPC and REST APIs
- Production experience with Kubernetes and containerized workloads
- Familiarity with observability tools such as Prometheus and Grafana
- GPU-based systems experience is preferred
- Experience with low-level hardware management such as BMCs or Redfish is preferred
- Experience operating large-scale distributed systems or high-throughput infrastructure is preferred
- Experience collaborating with or contributing to open-source projects such as Go or Redfish is preferred
Benefits
- Medical, dental, and vision insurance fully paid by CoreWeave
- Company-paid life insurance and voluntary supplemental life insurance
- Short- and long-term disability insurance
- Flexible Spending Account and Health Savings Account
- Tuition reimbursement
- Employee Stock Purchase Program participation
- Mental wellness benefits through Spring Health
- Family-forming support through Carrot
- Paid parental leave
- Flexible childcare support through Kinside
- 401(k) with employer match
- Flexible paid time off
- Catered lunch at office and data center locations
- Hybrid work is prioritized; remote work may be considered for candidates more than 30 miles from an office
- New hires attend onboarding at a company hub within their first month, and teams gather quarterly
Tech Stack
About CoreWeave
CoreWeave provides a GPU-accelerated cloud for AI training and inference, VFX, and rendering, with bare-metal instances, Kubernetes orchestration, and managed services to scale workloads. It sells on-demand and reserved capacity to AI labs, startups, and enterprises, and offers SaaS tools and hands-on support for deployment. Founded in 2017 and headquartered in New York, it is publicly traded on Nasdaq under the ticker CRWV.
