17 hours ago
Remote, India +5 moreSenior
Responsibilities
- Own end-to-end data extraction workflows across complex websites and deliver complete, accurate, structured datasets.
- Use custom workflows and available tools to accelerate data collection, validation, and task execution.
- Extract data from dynamic and interactive sources, including JavaScript-rendered content and changing site structures.
- Apply validation checks, cross-source consistency controls, formatting standards, and systematic verification to ensure data quality.
- Scale scraping operations through efficient batching or parallelization, monitor failures, and maintain stability against minor site changes.
- Handle complex data structures, anti-bot mechanisms, and large-scale web scraping workflows.
Requirements
- At least 5 years of relevant experience in data engineering, web scraping, automation, or software development.
- Strong Python web scraping experience with tools such as BeautifulSoup and Selenium, including dynamic content, AJAX, infinite scroll, and APIs via proxies.
- Ability to extract data from hierarchies, archived pages, inconsistent HTML, and other complex structures.
- Experience cleaning, normalizing, validating, and delivering structured datasets in CSV, JSON, or Google Sheets.
- Demonstrated experience handling anti-bot mechanisms and dynamic site structures at scale.
- Experience with AWS or equivalent cloud infrastructure and Docker in real workflows.
- Hands-on experience with LLM frameworks such as LangChain or OpenRouter for automation tasks.
- Strong attention to detail, independent troubleshooting ability, and a self-directed work ethic.
- Upper-intermediate English proficiency at B2 level or above.
- A GitHub link and a bachelor's or master's degree in engineering, applied mathematics, computer science, or a related technical field are pluses.
Benefits
- Fully remote freelance work.
- Part-time project with an estimated workload of 10–20 hours per week during active phases.
- Flexible project-based work through the Mindrift platform.
Tech Stack
Categories
Data Engineering
About Mindrift
Mindrift builds an expert-sourcing platform that connects domain specialists to project-based work training and evaluating generative AI models, including supervised fine-tuning, RLHF, evaluation, and red-teaming. It is built and operated by Toloka, part of Nebius Group, and run from Amsterdam, Netherlands. Work is fully remote and freelance, serving global technology companies developing and improving large AI systems.
