3 months ago
Remote, United StatesSenior
Responsibilities
- Own end-to-end data extraction workflows across complex websites and deliver complete, accurate, reliable structured datasets.
- Use tools and custom workflows to accelerate data collection, validation, and task execution.
- Extract data from dynamic and interactive websites, including JavaScript-rendered content, while adapting to changing site behavior.
- Apply validation checks, cross-source consistency controls, formatting requirements, and systematic verification to ensure data quality.
- Scale scraping operations through efficient batching or parallelization, monitor failures, and maintain stability against minor site structure changes.
Requirements
- At least 5 years of relevant experience in data engineering, web scraping, automation, or software development.
- Strong Python web scraping experience with tools such as BeautifulSoup and Selenium, including dynamic content, APIs, and proxies.
- Ability to extract data from complex structures such as hierarchies, archived pages, and inconsistent HTML.
- Experience with data cleaning, normalization, and validation, delivering structured datasets in formats such as CSV and JSON.
- Demonstrated experience handling anti-bot mechanisms and dynamic site structures at scale.
- Experience with AWS or equivalent cloud infrastructure and Docker containerization in real workflows.
- Hands-on experience with LLM frameworks such as LangChain or OpenRouter for automation tasks.
- Upper-intermediate English proficiency at B2 level or above.
- A relevant bachelor's or master's degree and a GitHub link are pluses.
Benefits
- Remote freelance position with an estimated 10–20 hours per week during active project phases.
- Part-time project opportunity with compensation of up to $45 per hour equivalent, depending on level and contribution pace.
Tech Stack
Categories
Data Engineering
