2 months ago
Remote, EMEASenior
Responsibilities
- Design and develop internal scheduling systems to maximize GPU utilization.
- Build APIs and tooling to manage GPU workload lifecycles and troubleshoot failures.
- Design and improve inference control planes to accelerate model deployment and inference request serving.
- Collaborate with research teams to continuously improve research velocity.
- Work with scalability and infrastructure teams to stabilize fault-tolerant training and keep GPU nodes healthy and fully utilized.
Requirements
- Strong programming skills in Go or similar languages.
- Strong systems engineering background in distributed systems, schedulers, control planes, or high-throughput data planes.
- Production experience with Kubernetes internals, including controllers, informers, and operators, rather than only deploying to Kubernetes.
- A focus on observability and debuggability when building production systems.
- Experience serving inference requests at large scale is preferred.
Benefits
- Fully remote work with flexible hours.
- 37 days per year of vacation and holidays.
- Health insurance allowance for the employee and dependents.
- 16 weeks of flexible, fully paid parental leave.
- Well-being, continuous learning, and home office allowances.
- Company-provided equipment.
- Frequent team get-togethers, including three-day monthly collaboration sessions in Paris and annual off-sites.
