2 months ago
Amsterdam, Netherlands or London, United KingdomSenior
Responsibilities
- Design and build the Serverless platform control plane, scheduler, runtime, autoscaler, and customer-facing APIs
- Solve cold-start latency, GPU scheduling under contention, multi-tenant isolation, fair-share quotas, and edge request routing
- Set technical direction, own architecture decisions, write design documentation, and align the team
- Raise engineering standards through code reviews, design reviews, and technical leadership
- Define SLOs, build observability, lead incident response, and improve the platform through postmortems
- Work directly with customers on architecture reviews, performance escalations, and production issues
- Partner with Product, GTM, and infrastructure teams to translate customer needs into a technical roadmap
Requirements
- 7+ years of professional software engineering experience shipping production distributed systems at scale
- Excellent knowledge of Golang or readiness to quickly switch to it
- Deep experience with Kubernetes and container orchestration
- Strong understanding of distributed-systems trade-offs and patterns including consistency, availability, queueing, backpressure, retries, idempotency, and multi-tenancy
- Experience designing and operating high-throughput, low-latency services
- Demonstrated ownership of difficult engineering problems and ability to drive design discussions and unblock teammates
- Ability to write reliable code and solve complex problems
- Teamwork-oriented approach
- Coding interview required
- Preferred experience with serverless or function-as-a-service platforms
- Preferred GPU scheduling experience, including Kubernetes device plugins, MIG, MPS, time-slicing, or NVIDIA GPU Operator
- Preferred ML inference experience with model loading and warm-pool strategies
- Preferred cold-start optimization at the runtime, image, or snapshot level
- Preferred experience writing Kubernetes operators with Go, controller-runtime, or kubebuilder
- Open-source contributions in serverless, scheduling, or inference ecosystems are a bonus
Benefits
- Hybrid work from the Amsterdam or London office
- Competitive compensation
- Career growth and learning opportunities
- Flexibility and ownership
- Collaborative and innovative culture
- Opportunity to work on impactful AI projects
- International environment and talented teams
Tech Stack
About Nebius
The Nebius AI Cloud brings powerful full-stack infrastructure for AI developers and practitioners across startups, enterprises and science institutes to build and deploy generative AI applications and rapidly deliver scientific breakthroughs by training and running ML models within a secure, high-performance, and cost-optimized cloud environment.
