3 months ago
Amsterdam, NetherlandsSenior
Responsibilities
- Onboard state-of-the-art open-source models into Nebius TokenFactory.
- Develop and optimize production-scale model-serving and inference systems.
- Implement cache-aware routing, NUMA-aware deployments, KV-cache offloading, disaggregated serving, and autoscaling with high-speed model loading.
- Maintain and extend forks of vLLM and TRT-LLM.
- Build performance, quality, smoke-testing, diagnostics, observability, traffic-replay, and automated rollout tooling.
- Develop hyperparameter optimization and automated search systems for inference and serverless deployment configurations.
- Collaborate with model builders, open-source communities, cloud teams, Solutions Architects, and hardware vendors.
Requirements
- Experience serving LLMs in production.
- Strong Python and/or Go programming skills.
- Experience designing and operating highly scalable, highly available distributed services.
- Contributions to vLLM, SGLang, TRT-LLM, or NVIDIA ecosystem open-source projects are beneficial.
- Deep understanding of KV-cache management, speculative decoding, and quantization is beneficial.
- Experience with LLM evaluation frameworks, performance benchmarking, and optimization is beneficial.
- Deep understanding of Kubernetes and familiarity with distributed serving architectures and autoscaling are beneficial.
- Knowledge of InfiniBand, RoCE, or high-performance networking is beneficial.
- Applicants must be authorized to work in the country where they apply and provide proof of employment eligibility.
Benefits
- Remote collaboration with in-person team meetings every one to two months at a Nebius location.
- Career growth and learning opportunities.
- Flexibility and ownership.
- Collaborative and innovative international environment.
- Opportunity to work on impactful AI projects.
Tech Stack
Categories
About Nebius
Nebius builds a full-stack AI cloud offering GPU compute, storage, and tools for training and deploying ML models for startups, enterprises, and research labs. It sells consumption-based cloud infrastructure (IaaS/PaaS) and managed services tailored to generative AI workloads, including large-scale model training and inference. The company is headquartered in Amsterdam and operates as an independent provider.
