4 months ago
Responsibilities
- Design and build multi-cloud orchestration systems with unified deployment and scheduling abstractions.
- Extend Kubernetes, including Dynamic Resource Allocation, for heterogeneous accelerators and multi-agent AI workflows.
- Implement intelligent load balancing and placement across cloud providers, regions, and hardware types.
- Build control-plane systems for allocating and managing heterogeneous accelerator capacity.
- Collaborate with Accelerator Systems Software engineers to expose low-level scheduling primitives.
- Support reliable operation of large-scale heterogeneous infrastructure through infrastructure-as-code and fleet-management practices.
Requirements
- Strong experience with Kubernetes internals, including custom controllers, schedulers, device plugins, CRDs, and the Dynamic Resource Allocation framework.
- Experience building or operating multi-cloud infrastructure with detailed knowledge of networking, storage, and compute differences across major providers.
- Familiarity with GPU and accelerator resource management in clusters, including MIG, time-slicing, device plugins, and topology-aware scheduling.
- Experience with infrastructure-as-code, fleet management, and reliability engineering for large-scale heterogeneous systems.
Benefits
- Competitive salary determined by skills and experience.
- Equity and ownership.
- Private healthcare.
- Visa sponsorship and relocation benefits.
- In-person work at the London office with provided tools, space, and setup.
