
Member of Technical Staff - Web Crawl Engineer
Reflection3 months ago
Responsibilities
- Build and operate web-scale crawling infrastructure that collects data across billions of URLs
- Design and optimize URL discovery, prioritization, scheduling, and crawl orchestration systems
- Develop distributed crawlers that acquire content while respecting site constraints and operational requirements
- Build content extraction, rendering, parsing, and normalization systems for diverse web formats
- Improve crawl coverage, freshness, efficiency, and quality through measurement and experimentation
- Design large-scale recrawling, change detection, and incremental update infrastructure
- Develop specialized crawlers for high-value domains, dynamic websites, and difficult-to-access sources
- Analyze crawl performance and web coverage to identify gaps and inefficiencies
- Build observability, monitoring, and reliability systems for crawl operations
- Debug production issues and improve the performance, scalability, and resilience of crawling infrastructure
Requirements
- Experience building large-scale web crawling, search indexing, content acquisition, or internet-scale data collection systems
- Strong understanding of crawling architectures, URL frontier management, scheduling, and distributed crawl coordination
- Experience with large-scale distributed systems using Ray, Spark, Beam, Flink, or similar frameworks
- Familiarity with content extraction, HTML parsing, browser automation, rendering systems, and modern web technologies
- Experience operating systems that process petabyte-scale datasets
- Strong systems engineering skills in reliability, observability, performance optimization, and debugging
- Experience designing experiments and using data to improve crawl quality, coverage, and efficiency
- Excellent communication skills and ability to reason clearly about system tradeoffs and operational constraints
- Experience building search engines, web indexes, or internet-scale crawling platforms is preferred
- Familiarity with anti-bot systems, dynamic web content, browser automation, and large-scale extraction pipelines is preferred
- Understanding of how web data is used in training and evaluating large language models is preferred
- Experience with distributed storage systems, content deduplication, and web-scale dataset management is preferred
Benefits
- Comprehensive medical, dental, vision, life, and disability insurance
- Fully paid parental leave for all new parents, including adoptive and surrogate journeys
- Financial support for family planning
- Paid time off
- Relocation support
- Daily provided lunch and dinner
- Regular off-sites and team celebrations
Tech Stack
Categories
BackendData Engineering
About Reflection
Reflection is a New York–based, privately held research lab developing open foundational AI models and agentic coding tools for developers, enterprises, and public-sector users. The team includes former researchers from DeepMind, OpenAI, and Anthropic, and their work focuses on transparent, customizable systems that organizations can deploy with ownership and control.