Clera

Scraping Engineer

Clera
Apply
1 day ago
Remote, WorldwideEntry Level

Responsibilities

  • Monitor and triage broken scrapers and data quality alerts.
  • Build and deploy web scraping scripts for sites requiring data extraction.
  • Validate scraped data, query data, and investigate discrepancies.
  • Debug and fix broken scrapers in production environments.
  • Create real-time dashboards for scraped data visualization and monitoring.
  • Collaborate with AI agents to address capability gaps, fix issues, and improve scripts.
  • Maintain reliable, high-throughput web extraction infrastructure processing millions of pages daily.

Requirements

  • At least 1 year of hands-on professional experience building or maintaining web scraping solutions.
  • Proficiency with TypeScript and Node.js.
  • Proficiency with SQL for data querying and validation.
  • Experience with Puppeteer or similar browser automation libraries.
  • Experience debugging and fixing broken scrapers in production.
  • Experience with RabbitMQ or similar message queues and asynchronous job processing systems.
  • Experience with Redis or similar in-memory caching systems.
  • Familiarity with Google Cloud Platform or equivalent cloud infrastructure.
  • Experience with bash scripting, Git, and gRPC.
  • Strong async communication skills and the ability to proactively flag blockers and provide updates.
  • Comfort with ambiguity and a fast-paced, early-stage environment.
  • BigQuery and dashboard or data visualization experience are pluses.

Benefits

  • Fully remote role open to candidates in any time zone.
  • Candidates must be available at least 8 hours daily with overlap during US business hours.
  • Visa sponsorship is not available.

Tech Stack

BashGitGoogle BigQueryGoogle Cloud PlatformgRPCNode.jsRabbitMQRedisSQLTypeScript

Categories

Data Engineering
Contact me