8 hours ago
London, United KingdomStaff+
Responsibilities
- Design and build runtime components for complex AI, simulation, and engineering workloads.
- Define workload, execution-environment, dependency, state, capability, failure, API, schema, and compatibility contracts across independently evolving systems.
- Build production runtime and control-plane software, adapters, integrations, conformance mechanisms, and validation systems for heterogeneous execution environments.
- Improve the reliability, observability, debuggability, and performance of distributed workload execution across large-scale GPU infrastructure.
- Design repeatable workload and benchmark environments and run experiments that distinguish real performance gains from scheduling behavior, warm-up effects, noise, and stragglers.
- Lead ambiguous systems problems, influence technical direction across teams, mentor engineers, and contribute to technical hiring and engineering standards.
Requirements
- Significant experience building complex systems software, distributed infrastructure, runtimes, workflow systems, or adjacent technology.
- Deep expertise in at least one of distributed systems, runtime systems, workflow or execution engines, programming languages/compilers/interpreters, cluster scheduling/orchestration, or high-performance and systems software.
- Strong software engineering fundamentals and production coding ability.
- Experience designing APIs, protocols, schemas, or contracts between independently evolving systems.
- Strong understanding of distributed-system failure modes, state, authority, retries, concurrency, and side effects.
- Strong technical judgment regarding abstraction, system boundaries, and operational complexity.
- Experience communicating and influencing architectural decisions across teams.
- Valuable additional experience includes Go, Rust, C/C++, Python, Kubernetes, containerized infrastructure, Argo, OSMO, Temporal, Ray, Kubeflow, GPU clusters, high-performance computing, simulation, robotics, autonomous systems, programming-language or compiler research, performance engineering, cloud infrastructure at scale, or platforms operated by autonomous software or AI agents.
Benefits
- Base salary range is £116,000 to £155,000, with discretionary bonus, equity awards, and comprehensive benefits based on eligibility.
- US-based full-time employee offerings include fully paid medical, dental, and vision insurance, life insurance, disability insurance, FSA, HSA, tuition reimbursement, ESPP participation, mental wellness benefits, family-forming support, paid parental leave, childcare support, and a 401(k) with employer match.
- Flexible PTO, catered lunch at office and data center locations, a casual work environment, and a culture focused on innovative disruption are offered.
About CoreWeave
CoreWeave provides a GPU-accelerated cloud for AI training and inference, VFX, and rendering, with bare-metal instances, Kubernetes orchestration, and managed services to scale workloads. It sells on-demand and reserved capacity to AI labs, startups, and enterprises, and offers SaaS tools and hands-on support for deployment. Founded in 2017 and headquartered in New York, it is publicly traded on Nasdaq under the ticker CRWV.
