22 hours ago
Remote, Chile +5 moreSenior
Responsibilities
- Own complete data extraction workflows across complex websites and deliver accurate, structured datasets.
- Use custom workflows and available tools to accelerate data collection, validation, and task execution.
- Extract data from dynamic and interactive sources, adapting to JavaScript-rendered content and changing site behavior.
- Apply validation, cross-source consistency checks, formatting controls, and systematic verification to maintain data quality.
- Scale scraping operations through batching or parallelization, monitor failures, and maintain stability against minor site changes.
Requirements
- At least 5 years of relevant experience in data engineering, web scraping, automation, or software development.
- Strong Python web-scraping experience with BeautifulSoup, Selenium or similar tools, including dynamic content and APIs via proxies.
- Experience extracting data from complex structures such as hierarchies, archived pages, and inconsistent HTML.
- Background in data cleaning, normalization, and validation using structured formats including CSV, JSON, and Google Sheets.
- Experience handling anti-bot mechanisms and dynamic site structures at scale.
- Experience with AWS or equivalent cloud infrastructure and Docker containerization.
- Hands-on experience with LLM frameworks such as LangChain or OpenRouter for automation tasks.
- Upper-intermediate English proficiency at B2 level or above.
- A GitHub link is a plus; a relevant bachelor's or master's degree is also a plus.
Benefits
- Remote, part-time freelance work with an estimated 10–20 hours per week during active project phases.
- Compensation of up to $25 per hour equivalent, depending on level and pace of contribution.
- Work on the Tendem project through Mindrift's specialist technology-project platform.
Tech Stack
Categories
Data Engineering
About Mindrift
Mindrift builds an expert-sourcing platform that connects domain specialists to project-based work training and evaluating generative AI models, including supervised fine-tuning, RLHF, evaluation, and red-teaming. It is built and operated by Toloka, part of Nebius Group, and run from Amsterdam, Netherlands. Work is fully remote and freelance, serving global technology companies developing and improving large AI systems.
