1 day ago
Remote, WorldwideEntry Level
Responsibilities
- Monitor and triage broken scrapers and data quality alerts.
- Build and deploy web scraping scripts for sites requiring data extraction.
- Validate scraped data, query data, and investigate discrepancies.
- Debug and fix broken scrapers in production environments.
- Create real-time dashboards for scraped data visualization and monitoring.
- Collaborate with AI agents to address capability gaps, fix issues, and improve scripts.
- Maintain reliable, high-throughput web extraction infrastructure processing millions of pages daily.
Requirements
- At least 1 year of hands-on professional experience building or maintaining web scraping solutions.
- Proficiency with TypeScript and Node.js.
- Proficiency with SQL for data querying and validation.
- Experience with Puppeteer or similar browser automation libraries.
- Experience debugging and fixing broken scrapers in production.
- Experience with RabbitMQ or similar message queues and asynchronous job processing systems.
- Experience with Redis or similar in-memory caching systems.
- Familiarity with Google Cloud Platform or equivalent cloud infrastructure.
- Experience with bash scripting, Git, and gRPC.
- Strong async communication skills and the ability to proactively flag blockers and provide updates.
- Comfort with ambiguity and a fast-paced, early-stage environment.
- BigQuery and dashboard or data visualization experience are pluses.
Benefits
- Fully remote role open to candidates in any time zone.
- Candidates must be available at least 8 hours daily with overlap during US business hours.
- Visa sponsorship is not available.
Tech Stack
Categories
Data Engineering
