4 days ago
Remote, Mexico +5 moreSenior
Responsibilities
- Own end-to-end data extraction workflows across complex websites and deliver complete, accurate, structured datasets.
- Use custom workflows and tools to accelerate data collection, validation, and task execution.
- Extract data from dynamic and interactive websites, including JavaScript-rendered content, AJAX, infinite scroll, and changing site structures.
- Apply data-quality controls including validation, cross-source consistency checks, formatting compliance, and systematic verification.
- Scale scraping operations through efficient batching or parallelization, monitor failures, and maintain stability against minor site changes.
Requirements
- At least 5 years of relevant experience in data engineering, web scraping, automation, or software development.
- Strong Python web-scraping experience with tools such as BeautifulSoup and Selenium, including dynamic content and APIs via proxies.
- Ability to extract data from complex structures such as hierarchies, archived pages, and inconsistent HTML.
- Experience with data cleaning, normalization, validation, and delivery of structured CSV, JSON, or Google Sheets datasets.
- Experience handling anti-bot mechanisms and dynamic site structures at scale.
- Experience with AWS or equivalent cloud infrastructure and Docker containerization.
- Hands-on experience with LLM frameworks such as LangChain or OpenRouter for automation tasks.
- English proficiency at Upper-intermediate (B2) level or above.
- A GitHub link is a plus, along with a methodical, detail-oriented, self-directed approach.
Benefits
- Remote freelance work based in Mexico City, with an estimated 10–20 hours per week during active project phases.
- Part-time project participation with workload dependent on project requirements and active phases.
Tech Stack
Categories
Data Engineering
