13 hours ago
Remote, EMEASenior
Responsibilities
- Design and implement model routers that select models or reasoning profiles based on quality, cost, latency, reliability, and related objectives.
- Develop provider and protocol abstractions, versioned schemas, execution contracts, and portable agent skill and memory systems.
- Create benchmark suites and evaluation protocols for models, routers, memories, skills, harnesses, and agent workflows.
- Research retrieval, context selection and compaction, identity, provenance, distillation, self-improving harnesses, multi-agent learning, and automated skill creation.
- Write robust research software, APIs, integration layers, distributed systems, and test infrastructure for reproducible experimentation.
- Collaborate across research and engineering teams and communicate findings through technical reports, demonstrations, open-source releases, benchmarks, and publications.
Requirements
- Profound understanding of machine learning, large language models, or statistical decision-making.
- Deep expertise in at least one relevant area such as model routing, recommender systems, agent systems, retrieval and memory, model evaluation, distributed systems, or protocol and API design.
- Experience building and evaluating modern language-model or agentic systems with tool use and multi-turn workflows.
- Experience designing, executing, and analyzing machine-learning experiments with statistical rigor, including held-out testing, out-of-domain evaluation, uncertainty, and reproducibility.
- Strong software-engineering and algorithm-design skills, including excellent Python proficiency and experience with APIs, data schemas, distributed services, testing, observability, code review, and CI/CD.
- Ability to reason about security, privacy, provenance, permissions, failure modes, and user control in agent systems.
- Strong communication and technical leadership skills across research and engineering disciplines.
- Preferred experience with model routers, cascades, mixture-of-experts, recommenders, cost-aware inference, multiple model providers, agent harnesses, retrieval systems, vector search, knowledge graphs, benchmark suites, distillation, reinforcement learning, preference learning, reward modeling, or automated skill generation.
- Preferred proficiency in TypeScript, Go, Rust, or another systems language in addition to Python.
- Preferred experience with secure authentication, sandboxing, privacy-preserving telemetry, policy-enforced execution, distributed data-processing, model-training, or inference systems.
- A PhD in Computer Science, Machine Learning, Artificial Intelligence, or a related technical field, or equivalent practical experience, is preferred.
- A track record of impactful publications, open-source contributions, deployed AI systems, or delivered products and research prototypes is preferred.
Benefits
- Competitive compensation.
- Career growth and learning opportunities.
- Flexibility and ownership.
- Collaborative and innovative culture.
- Opportunity to work on impactful AI projects.
- International environment with talented teams.
Tech Stack
Categories
AI ResearchML Engineering
About Nebius
Nebius builds a full-stack AI cloud offering GPU compute, storage, and tools for training and deploying ML models for startups, enterprises, and research labs. It sells consumption-based cloud infrastructure (IaaS/PaaS) and managed services tailored to generative AI workloads, including large-scale model training and inference. The company is headquartered in Amsterdam and operates as an independent provider.
