2 months ago
Base Salary
$182k - $242k/yr
Responsibilities
- Own the reliability and performance of a Kubernetes-based data platform end to end
- Architect and operate highly available, multi-region systems meeting stringent uptime and latency standards
- Scale infrastructure and optimize deployment pipelines for production-grade services
- Strengthen the platform’s security posture through cloud-native security practices
- Build, test, and manage scalable systems supporting data ingestion, transformation, analytics, and internal AI workloads
- Develop and operate distributed systems, databases, backend APIs, and internal platform capabilities
- Build and maintain observability covering metrics, logging, and distributed tracing
- Lead production ownership activities including incident response, reliability objectives, error budgets, and postmortems
- Perform performance tuning, capacity planning, resource optimization, and automated environment provisioning
Requirements
- 7+ years of hands-on experience in platform engineering, infrastructure engineering, or highly scalable distributed systems
- Strong foundation in software design, development, and algorithmic problem-solving
- Experience operating data platforms or data-intensive workloads with distributed processing and streaming frameworks such as Spark, Airflow, Kafka, or Flink
- Deep expertise in Kubernetes cluster design, operations, and production troubleshooting across containerized environments
- Experience building and operating CI/CD pipelines using Argo CD and GitHub Actions
- Experience owning mission-critical systems with high availability requirements of at least 99.99% uptime, including incident response, SLI/SLO/SLA definition, error budget management, and blameless postmortems
- Hands-on experience developing large-scale distributed systems, databases, and backend APIs
- Experience designing and operating geo-replicated, active-active, multi-region systems, including traffic routing, failover, and data consistency tradeoffs
- Experience building full-stack observability solutions for metrics, logging, and distributed tracing using tools such as Prometheus, Grafana, and OpenTelemetry
- Proficiency with infrastructure-as-code tools such as Helm, Terraform, or Pulumi and automated environment provisioning
- Strong command of performance tuning, capacity planning, and resource optimization in distributed environments
- Hands-on experience with cloud-native security practices including secrets management, network policies, and vulnerability scanning
- Preferred experience in compliance-driven environments and knowledge of GDPR, SOC 2, HIPAA, or SOX
- Preferred experience designing internal developer platforms or self-service infrastructure tooling
- Must meet applicable US export-control eligibility requirements for access to export-controlled information
Benefits
- Medical, dental, and vision insurance fully paid by CoreWeave for US-based employees
- Company-paid life insurance, voluntary supplemental life insurance, and short- and long-term disability insurance
- Flexible Spending Account and Health Savings Account
- Tuition reimbursement
- Employee Stock Purchase Program eligibility
- Mental wellness benefits through Spring Health
- Family-forming support through Carrot
- Paid parental leave and flexible childcare support through Kinside
- 401(k) with an employer match
- Flexible paid time off
- Catered lunch at office and data center locations
- Casual work environment and culture focused on innovative disruption
Tech Stack
About CoreWeave
CoreWeave is the Essential Cloud for AI. CoreWeave is a cloud purpose-built for scaling, supporting, and accelerating GenAI. We’re a comprehensive platform and strategic partner designed to tackle today—and tomorrow’s—challenges of deploying AI at scale. We manage the complexities of AI growth to make supercomputing accessible and push the limits of what’s possible. Our teams create modern solutions to support modern technology. Get the premier choice for working with GenAI workloads.
