27 days ago
Base Salary
$165k - $242k/yr
Responsibilities
- Build and ship production machine learning systems for infrastructure observability, troubleshooting, and optimization.
- Develop anomaly detection, time-series analysis, event correlation, ranking, recommendation, classification, and root-cause inference approaches.
- Create datasets, experiments, and evaluation frameworks for model quality, robustness, usefulness, and failure analysis.
- Build data pipelines, feature-generation workflows, inference services, and feedback loops.
- Integrate ML capabilities with Mission Control, Grafana, and related observability services.
- Monitor and improve the quality, latency, reliability, and cost of ML-powered production services.
- Investigate data and model failures and implement durable fixes.
- Contribute to technical designs, code reviews, testing standards, operational practices, documentation, mentorship, and cross-functional delivery.
Requirements
- Several years of experience designing and shipping machine learning systems in production.
- Strong Python software engineering skills; Go or another systems-oriented language is a plus.
- Solid understanding of model selection, feature engineering, experimentation, evaluation, and failure analysis.
- Experience with time-series, event, log, metric, trace, or other operational data.
- Experience building reliable data and inference services rather than only notebooks or offline prototypes.
- Experience defining evaluation methodologies for ambiguous or domain-specific ML problems.
- Strong debugging and systems-thinking skills involving data quality, distributed systems, latency, and operational failure modes.
- Experience taking an ML capability from initial hypothesis through production launch and iteration.
- Excellent communication and collaboration skills across engineering, research, product, and infrastructure teams.
- Track record of owning complex technical work and delivering results with appropriate guidance and autonomy.
- Preferred experience with Grafana, Prometheus, VictoriaMetrics, ClickHouse, Loki, Kafka, Kubernetes, cloud infrastructure, anomaly detection, incident intelligence, search, recommendations, ranking, LLM evaluation, post-training, retrieval, tool use, grounded generation, human-in-the-loop workflows, access control, auditability, or operational safety requirements.
Benefits
- Base salary range of $165,000 to $242,000, plus eligibility for discretionary bonus, equity awards, and a comprehensive benefits program.
- Medical, dental, and vision insurance fully paid by CoreWeave, plus company-paid life insurance and voluntary supplemental life insurance.
- Short- and long-term disability insurance, Flexible Spending Account, and Health Savings Account.
- Tuition reimbursement and participation in the Employee Stock Purchase Program.
- Mental wellness benefits through Spring Health and family-forming support through Carrot.
- Paid parental leave and flexible, full-service childcare support through Kinside.
- 401(k) with an employer match and flexible paid time off.
- Catered lunch at office and data center locations, a casual work environment, and an innovation-focused culture.
Tech Stack
Categories
About CoreWeave
CoreWeave provides a GPU-accelerated cloud for AI training and inference, VFX, and rendering, with bare-metal instances, Kubernetes orchestration, and managed services to scale workloads. It sells on-demand and reserved capacity to AI labs, startups, and enterprises, and offers SaaS tools and hands-on support for deployment. Founded in 2017 and headquartered in New York, it is publicly traded on Nasdaq under the ticker CRWV.
