CoreWeave

Staff Software Engineer, MetalDev

CoreWeave
Apply
2 months ago
Sunnyvale, CA, USA or New York, NY, USAStaff+
H1B sponsor

Base Salary

$207k - $275k/yr

Responsibilities

  • Design, build, and operate Go-based services managing the lifecycle of large-scale GPU data center infrastructure.
  • Develop automation for data center bring-up, hardware discovery, health monitoring, remediation, and production operations.
  • Build APIs, services, and workflows for BMCs, firmware state, server health, and rack-level infrastructure.
  • Improve observability, alerting, and operational tooling for rapid detection and resolution of production issues.
  • Translate incidents and hardware failure modes into software improvements that increase platform resilience.
  • Partner with hardware-adjacent, infrastructure, operations, and software teams on safe fleet-scale systems.
  • Provide technical leadership through design reviews, code reviews, architectural guidance, mentorship, and technical project leadership.
  • Lead incident response and postmortems while making architecture decisions that balance reliability, simplicity, scalability, and operational burden.

Requirements

  • Bachelor’s, master’s, or PhD in Computer Science or a related field, or equivalent experience.
  • At least 8 years of software engineering experience focused on infrastructure, cloud engineering, and distributed databases, particularly in large-scale data center and cloud environments.
  • Expertise in Go and experience building REST/gRPC APIs for mission-critical platforms.
  • Strong experience architecting and scaling cloud-native Kubernetes infrastructure and distributed services.
  • Demonstrated ability to mentor engineers, lead technical projects, and influence engineering strategy across teams.
  • Experience contributing to and collaborating with open source communities.
  • Ability to apply data-driven methods to reliability, optimization, and continuous improvement.
  • Strong communication skills with technical and non-technical stakeholders.
  • Hands-on experience with Prometheus, Grafana, PromQL, CI/CD pipelines, and large fleets of GPU servers.
  • Experience leading incident response, postmortems, and service reliability improvements.
  • Working knowledge of Kafka, ClickHouse, CRDB, DMTF, RedFish APIs, and GPU servers is preferred.

Benefits

  • Medical, dental, and vision insurance fully paid by CoreWeave for US-based full-time employees.
  • Company-paid life insurance, voluntary supplemental life insurance, and short- and long-term disability insurance.
  • Flexible Spending Account and Health Savings Account.
  • Tuition reimbursement and participation in the Employee Stock Purchase Program.
  • Mental wellness benefits through Spring Health and family-forming support through Carrot.
  • Paid parental leave and flexible full-service childcare support through Kinside.
  • 401(k) with an employer match.
  • Flexible paid time off.
  • Catered lunch each day in office and data center locations.
  • Casual work environment and a culture focused on innovative disruption.
  • Benefits vary by location and are shared during the hiring process for non-US roles.

Tech Stack

Apache KafkaClickHouseGoGrafanagRPCKubernetesPrometheus

Categories

BackendDevOpsSite Reliability
CoreWeave

About CoreWeave

1,001-5,000 employees

CoreWeave provides a GPU-accelerated cloud for AI training and inference, VFX, and rendering, with bare-metal instances, Kubernetes orchestration, and managed services to scale workloads. It sells on-demand and reserved capacity to AI labs, startups, and enterprises, and offers SaaS tools and hands-on support for deployment. Founded in 2017 and headquartered in New York, it is publicly traded on Nasdaq under the ticker CRWV.

Contact me