12 hours ago
Zürich, SwitzerlandSenior / Staff+
Responsibilities
- Design, implement, and operate core runtime services that serve search queries at scale.
- Build and optimize query processing, retrieval orchestration, and response assembly flows under strict latency budgets.
- Improve CPU, memory, networking, data access, throughput, latency, reliability, and cost efficiency in production services.
- Define observability primitives including structured logs, metrics, traces, and latency breakdowns.
- Collaborate with indexing, ML, and product teams to integrate retrieval and ranking components while keeping ML logic decoupled from core systems.
- Support experimentation through controlled rollouts, rigorous benchmarking, and well-tested systems.
Requirements
- 5+ years of experience as a software engineer working on production backend systems.
- Strong hands-on expertise with C++ or Rust in real-world, high-load services.
- Experience building and operating low-latency, user-facing systems handling thousands of requests per second under strict latency constraints.
- Systems-level understanding of CPU, memory, networking, and data access performance.
- Experience deploying, debugging, operating, and rolling back production code.
- Strong end-to-end thinking about request flows and the ability to make pragmatic tradeoffs among correctness, latency, and development velocity.
- Effective cross-functional collaboration with engineering, ML, and product teams.
- Preferred experience includes DBMS internals, cloud infrastructure, high-load web applications or large-scale APIs, performance-critical systems, low-level performance tuning, open-source contributions, competitive programming, CTF participation, SHAD or similar programs, conference talks, or technical publications.
- Coding interviews are part of the hiring process.
Benefits
- Competitive compensation.
- Career growth and learning opportunities.
- Flexibility and ownership.
- Collaborative and innovative culture.
- Opportunity to work on impactful AI projects.
- International environment and talented teams.
- Applicants must be authorized to work in the country in which they apply and must provide proof of employment eligibility at hire.
About Nebius
Nebius builds a full-stack AI cloud offering GPU compute, storage, and tools for training and deploying ML models for startups, enterprises, and research labs. It sells consumption-based cloud infrastructure (IaaS/PaaS) and managed services tailored to generative AI workloads, including large-scale model training and inference. The company is headquartered in Amsterdam and operates as an independent provider.
