3 hours ago
Base Salary
$182k - $242k/yr
Responsibilities
- Lead incident response, root cause analysis, post-incident reviews, incident documentation, and long-term service improvements.
- Develop and continuously improve incident response playbooks and communicate clearly with stakeholders during incidents.
- Improve core service performance, robustness, supportability, resilience, and disaster recovery at scale.
- Own observability and health monitoring using Prometheus and Grafana, including dashboards, alerts, and proactive bottleneck detection.
- Lead automation for incident detection and recovery and design solutions that improve operational efficiency and stability.
- Define and drive incident-management KPIs and SLAs aligned with organizational reliability objectives.
- Create CI/CD pipelines and document hardware automation workflows and processes.
- Support the full server hardware lifecycle from provisioning through end-of-life by troubleshooting bugs and automating common tasks.
- Partner with Fleet Operations to build scalable self-service tooling and reduce escalation overhead.
- Participate in on-call rotation and continuously reduce support queries and production incidents.
Requirements
- 7+ years of experience in cloud operations, site reliability engineering, or related technical roles.
- Understanding of cloud platforms including Kubernetes, AWS, and GCP, plus basic cloud infrastructure knowledge.
- Familiarity with incident management practices and frameworks, including ITIL and SRE best practices.
- Proficiency with Go or Python.
- Prior experience with Prometheus and Grafana.
- Experience deploying containerized applications using Kubernetes.
- Experience serving on an on-call rotation supporting production services.
- Strong documentation, analytical, problem-solving, and attention-to-detail skills.
Benefits
- Base salary range is $182,000 to $242,000, plus eligibility for a discretionary bonus, equity awards, and a comprehensive benefits program.
- Medical, dental, and vision insurance are 100% paid by CoreWeave, along with company-paid life insurance.
- Additional benefits include disability insurance, Flexible Spending Account, Health Savings Account, tuition reimbursement, ESPP participation, mental wellness benefits, family-forming support, paid parental leave, childcare support, and a 401(k) with employer match.
- Flexible PTO, a casual work environment, and catered lunch at office and data center locations are provided.
- The role is a full-time US position; benefits vary for roles in other locations and are shared during the hiring process.
- The position requires access to export-controlled information and is subject to applicable US export-control eligibility requirements.
Tech Stack
Categories
DevOpsSite Reliability
About CoreWeave
CoreWeave is the Essential Cloud for AI. CoreWeave is a cloud purpose-built for scaling, supporting, and accelerating GenAI. We’re a comprehensive platform and strategic partner designed to tackle today—and tomorrow’s—challenges of deploying AI at scale. We manage the complexities of AI growth to make supercomputing accessible and push the limits of what’s possible. Our teams create modern solutions to support modern technology. Get the premier choice for working with GenAI workloads.
