4 hours ago
Base Salary
$207k - $275k/yr
Responsibilities
- Design, build, and operate Go-based services managing the lifecycle of large-scale GPU data center infrastructure.
- Develop automation for data center bring-up, hardware discovery, health monitoring, remediation, and production operations.
- Build APIs, services, and workflows for BMCs, firmware state, server health, and rack-level infrastructure.
- Improve observability, alerting, and operational tooling for rapid detection and resolution of production issues.
- Translate incidents and hardware failure modes into software improvements that increase platform resilience.
- Partner with hardware-adjacent, infrastructure, operations, and software teams on safe fleet-scale systems.
- Provide technical leadership through design reviews, code reviews, architectural guidance, mentorship, and technical project leadership.
- Lead incident response and postmortems while making architecture decisions that balance reliability, simplicity, scalability, and operational burden.
Requirements
- Bachelor’s, master’s, or PhD in Computer Science or a related field, or equivalent experience.
- At least 8 years of software engineering experience focused on infrastructure, cloud engineering, and distributed databases, particularly in large-scale data center and cloud environments.
- Expertise in Go and experience building REST/gRPC APIs for mission-critical platforms.
- Strong experience architecting and scaling cloud-native Kubernetes infrastructure and distributed services.
- Demonstrated ability to mentor engineers, lead technical projects, and influence engineering strategy across teams.
- Experience contributing to and collaborating with open source communities.
- Ability to apply data-driven methods to reliability, optimization, and continuous improvement.
- Strong communication skills with technical and non-technical stakeholders.
- Hands-on experience with Prometheus, Grafana, PromQL, CI/CD pipelines, and large fleets of GPU servers.
- Experience leading incident response, postmortems, and service reliability improvements.
- Working knowledge of Kafka, ClickHouse, CRDB, DMTF, RedFish APIs, and GPU servers is preferred.
Benefits
- Medical, dental, and vision insurance fully paid by CoreWeave for US-based full-time employees.
- Company-paid life insurance, voluntary supplemental life insurance, and short- and long-term disability insurance.
- Flexible Spending Account and Health Savings Account.
- Tuition reimbursement and participation in the Employee Stock Purchase Program.
- Mental wellness benefits through Spring Health and family-forming support through Carrot.
- Paid parental leave and flexible full-service childcare support through Kinside.
- 401(k) with an employer match.
- Flexible paid time off.
- Catered lunch each day in office and data center locations.
- Casual work environment and a culture focused on innovative disruption.
- Benefits vary by location and are shared during the hiring process for non-US roles.
Tech Stack
About CoreWeave
CoreWeave is the Essential Cloud for AI. CoreWeave is a cloud purpose-built for scaling, supporting, and accelerating GenAI. We’re a comprehensive platform and strategic partner designed to tackle today—and tomorrow’s—challenges of deploying AI at scale. We manage the complexities of AI growth to make supercomputing accessible and push the limits of what’s possible. Our teams create modern solutions to support modern technology. Get the premier choice for working with GenAI workloads.
