4 days ago
Remote, BrazilSenior
Responsibilities
- Own end-to-end data extraction workflows across complex websites and deliver accurate, structured datasets.
- Use tools and custom workflows to accelerate data collection, validation, and task execution.
- Adapt scraping approaches for JavaScript-rendered content, dynamic sources, APIs, and changing site behavior.
- Apply validation, cross-source consistency checks, formatting controls, and systematic verification to maintain data quality.
- Scale scraping operations through batching or parallelization, monitor failures, and maintain stability against site changes.
Requirements
- At least 5 years of relevant experience in data engineering, web scraping, automation, or software development.
- Strong experience with Python web scraping using BeautifulSoup, Selenium, or similar tools, including dynamic content and APIs via proxies.
- Ability to extract data from complex structures such as hierarchies, archived pages, and inconsistent HTML.
- Experience with data cleaning, normalization, validation, and structured outputs including CSV, JSON, and Google Sheets.
- Experience handling anti-bot mechanisms and dynamic site structures at scale.
- Experience with AWS or equivalent cloud infrastructure and Docker containerization in real workflows.
- Hands-on experience with LLM frameworks such as LangChain, OpenRouter, or similar for automation tasks.
- Upper-intermediate English proficiency at B2 level or above.
- A bachelor’s or master’s degree in engineering, applied mathematics, computer science, or a related technical field is a plus; a GitHub link is also a plus.
- Strong attention to detail, data accuracy, independent troubleshooting, and self-directed work habits.
Benefits
- Remote freelance work based in Brazil.
- Part-time project workload estimated at 10–20 hours per week during active phases, with no guaranteed workload.
- Compensation of up to $25 per hour equivalent, depending on level and pace of contribution.
Tech Stack
Categories
Data Engineering
