Reflection

Member of Technical Staff - Web Crawl Engineer

Reflection
Apply
3 months ago

Responsibilities

  • Build and operate web-scale crawling infrastructure that collects data across billions of URLs
  • Design and optimize URL discovery, prioritization, scheduling, and crawl orchestration systems
  • Develop distributed crawlers that acquire content while respecting site constraints and operational requirements
  • Build content extraction, rendering, parsing, and normalization systems for diverse web formats
  • Improve crawl coverage, freshness, efficiency, and quality through measurement and experimentation
  • Design large-scale recrawling, change detection, and incremental update infrastructure
  • Develop specialized crawlers for high-value domains, dynamic websites, and difficult-to-access sources
  • Analyze crawl performance and web coverage to identify gaps and inefficiencies
  • Build observability, monitoring, and reliability systems for crawl operations
  • Debug production issues and improve the performance, scalability, and resilience of crawling infrastructure

Requirements

  • Experience building large-scale web crawling, search indexing, content acquisition, or internet-scale data collection systems
  • Strong understanding of crawling architectures, URL frontier management, scheduling, and distributed crawl coordination
  • Experience with large-scale distributed systems using Ray, Spark, Beam, Flink, or similar frameworks
  • Familiarity with content extraction, HTML parsing, browser automation, rendering systems, and modern web technologies
  • Experience operating systems that process petabyte-scale datasets
  • Strong systems engineering skills in reliability, observability, performance optimization, and debugging
  • Experience designing experiments and using data to improve crawl quality, coverage, and efficiency
  • Excellent communication skills and ability to reason clearly about system tradeoffs and operational constraints
  • Experience building search engines, web indexes, or internet-scale crawling platforms is preferred
  • Familiarity with anti-bot systems, dynamic web content, browser automation, and large-scale extraction pipelines is preferred
  • Understanding of how web data is used in training and evaluating large language models is preferred
  • Experience with distributed storage systems, content deduplication, and web-scale dataset management is preferred

Benefits

  • Comprehensive medical, dental, vision, life, and disability insurance
  • Fully paid parental leave for all new parents, including adoptive and surrogate journeys
  • Financial support for family planning
  • Paid time off
  • Relocation support
  • Daily provided lunch and dinner
  • Regular off-sites and team celebrations

Tech Stack

Apache BeamApache FlinkApache SparkHTML

Categories

BackendData Engineering
Reflection

About Reflection

201-500 employees

Reflection is a New York–based, privately held research lab developing open foundational AI models and agentic coding tools for developers, enterprises, and public-sector users. The team includes former researchers from DeepMind, OpenAI, and Anthropic, and their work focuses on transparent, customizable systems that organizations can deploy with ownership and control.

Contact me