12 hours ago
Zürich, SwitzerlandSenior / Staff+
Responsibilities
- Design, implement, and operate web-scale crawling systems for acquiring internet content.
- Build ingestion workflows for crawlers, structured feeds, and partner integrations.
- Develop crawl scheduling, prioritization, recrawl policies, and freshness strategies.
- Build URL discovery, deduplication, content extraction, and crawl orchestration systems.
- Ensure reliable operation of high-throughput crawling infrastructure.
- Define observability and quality metrics for crawl coverage, freshness, throughput, and content quality.
- Monitor resource usage, bandwidth consumption, and infrastructure cost.
- Collaborate with indexing and ML teams to meet retrieval and ranking requirements.
- Enable safe experimentation with crawling strategies and content acquisition policies.
Requirements
- 5+ years of experience building backend or distributed systems.
- Strong expertise in Go or C++.
- Experience with large-scale distributed systems handling 10k+ RPS, billions of URLs, or high-throughput pipelines.
- Understanding of HTTP, DNS, TLS, crawling, scraping, and content extraction.
- Experience operating production systems and debugging failures in distributed environments.
- Strong understanding of scalability, fault tolerance, and resource management.
- Experience with web crawling, streaming data pipelines, event-driven systems, messaging platforms, distributed schedulers, queues, asynchronous processing, or large-scale data-processing frameworks is advantageous.
- Experience in ad tech, social networks, search engines, or other large-scale content platforms is advantageous.
- Willingness to complete coding interviews.
Benefits
- Competitive compensation.
- Career growth and learning opportunities.
- Flexibility and ownership.
- Collaborative and innovative culture.
- Opportunity to work on impactful AI projects.
- International environment and talented teams.
Tech Stack
Categories
BackendData Engineering
About Nebius
Nebius builds a full-stack AI cloud offering GPU compute, storage, and tools for training and deploying ML models for startups, enterprises, and research labs. It sells consumption-based cloud infrastructure (IaaS/PaaS) and managed services tailored to generative AI workloads, including large-scale model training and inference. The company is headquartered in Amsterdam and operates as an independent provider.
