4 days ago
Remote, Brazil +5 moreSenior
Responsibilities
- Own end-to-end data extraction workflows across complex websites and deliver accurate, structured datasets.
- Use custom workflows and tools to collect, validate, and process data efficiently.
- Extract data from dynamic and interactive websites, including JavaScript-rendered content, while adapting to site changes.
- Apply data-quality checks, cross-source consistency controls, formatting standards, and systematic verification.
- Scale scraping operations through batching or parallelization, monitor failures, and maintain workflow stability.
Requirements
- At least 5 years of relevant experience in data engineering, web scraping, automation, or software development.
- Strong Python web scraping experience with tools such as BeautifulSoup and Selenium, including dynamic content and APIs via proxies.
- Ability to extract data from complex structures such as hierarchies, archived pages, and inconsistent HTML.
- Experience with data cleaning, normalization, validation, and structured outputs such as CSV, JSON, and Google Sheets.
- Experience handling anti-bot mechanisms and changing website structures at scale.
- Experience with AWS or equivalent cloud infrastructure and Docker containerization.
- Hands-on experience with LLM frameworks such as LangChain or OpenRouter for automation tasks.
- Upper-intermediate English proficiency (B2) or above.
- A bachelor’s or master’s degree in engineering, applied mathematics, computer science, or a related technical field is a plus; a GitHub link is also a plus.
Benefits
- Remote freelance role with an estimated workload of 10–20 hours per week during active project phases.
- Compensation of up to $25 per hour equivalent, depending on contribution level and pace.
Tech Stack
Categories
Data Engineering
