5 hours ago
Base Salary
$188k - $250k/yr
Responsibilities
- Lead the technical design and delivery of production machine learning systems for infrastructure observability, troubleshooting, and optimization.
- Develop anomaly detection, time-series analysis, event correlation, ranking, recommendation, classification, and root-cause inference approaches.
- Build evaluation frameworks and datasets for accuracy, usefulness, robustness, and safety in operational scenarios.
- Design data pipelines, feature-generation workflows, model-serving paths, feedback loops, and model lifecycle practices.
- Integrate ML capabilities with telemetry platforms, Mission Control, Grafana, and related observability services.
- Establish standards for experimentation, evaluation, monitoring, reproducibility, and operational ownership.
- Balance model quality, latency, cost, interpretability, reliability, and ease of operation.
- Partner across engineering, research, product, infrastructure, and customer-facing teams.
- Mentor engineers and provide technical leadership across organizational boundaries.
Requirements
- Significant experience designing and shipping machine learning systems in production.
- Strong Python software engineering skills; Go or another systems-oriented language is a plus.
- Deep knowledge of machine learning fundamentals, model selection, feature engineering, experimentation, evaluation, and failure analysis.
- Experience working with time-series, event, log, metric, trace, or other operational data at meaningful scale.
- Ability to define evaluation methodologies for ambiguous or domain-specific ML problems.
- Strong systems thinking involving distributed systems, data quality, latency, observability, and operational failure modes.
- Excellent communication and collaboration skills across research, product, infrastructure, and customer-facing teams.
- Track record of technical leadership, mentorship, and cross-organizational influence.
- Experience with observability platforms or technologies such as Grafana, Prometheus, VictoriaMetrics, ClickHouse, Loki, or Kafka is preferred.
- Experience with Kubernetes and cloud infrastructure, anomaly detection, incident intelligence, search, recommendations, ranking, or ML products for technical users is preferred.
- Experience with large language model evaluation, post-training, retrieval, tool use, grounded generation, human-in-the-loop workflows, access control, auditability, or operational safety is preferred.
Benefits
- Base salary range of $188,000 to $250,000, plus potential discretionary bonus and equity awards.
- Medical, dental, and vision insurance fully paid by CoreWeave for eligible US-based full-time employees.
- Company-paid life insurance, supplemental life insurance, short- and long-term disability insurance, FSA, and HSA.
- Tuition reimbursement and eligibility to participate in the Employee Stock Purchase Program.
- Mental wellness benefits, family-forming support, paid parental leave, and childcare support.
- 401(k) with employer match and flexible paid time off.
- Catered lunch at office and data center locations and a casual work environment.
- Benefits vary by location; the role requires compliance with applicable export-control access requirements.
Tech Stack
Categories
About CoreWeave
CoreWeave provides a GPU-accelerated cloud for AI training and inference, VFX, and rendering, with bare-metal instances, Kubernetes orchestration, and managed services to scale workloads. It sells on-demand and reserved capacity to AI labs, startups, and enterprises, and offers SaaS tools and hands-on support for deployment. Founded in 2017 and headquartered in New York, it is publicly traded on Nasdaq under the ticker CRWV.
