2 days ago
Remote, India +5 moreSenior
Responsibilities
- Own end-to-end scraping and data extraction workflows across complex websites.
- Collect, clean, normalize, validate, and deliver structured datasets.
- Handle JavaScript-rendered content, AJAX, infinite scrolling, APIs, proxies, anti-bot mechanisms, and changing site structures.
- Scale scraping operations through batching or parallelization while monitoring failures and maintaining stability.
- Apply cross-source consistency checks, formatting standards, and systematic verification before delivery.
- Troubleshoot independently and use available tools and custom workflows to accelerate task execution.
Requirements
- At least 5 years of relevant experience in data engineering, web scraping, automation, or software development.
- Strong Python web-scraping experience with BeautifulSoup, Selenium or similar tools, including dynamic content.
- Ability to extract data from complex hierarchies, archived pages, and inconsistent HTML structures.
- Experience with data cleaning, normalization, validation, and structured outputs such as CSV and JSON.
- Experience handling anti-bot mechanisms and dynamic site structures at scale.
- Experience with AWS or equivalent cloud infrastructure and Docker in real workflows.
- Hands-on experience with LLM frameworks such as LangChain, OpenRouter, or similar for automation.
- Upper-intermediate English proficiency at B2 level or above.
- A bachelor’s or master’s degree in engineering, applied mathematics, computer science, or a related technical field is preferred, not required.
- A GitHub link is a plus.
Benefits
- Remote freelance and part-time arrangement.
- Estimated workload of 10–20 hours per week during active project phases, with no guaranteed workload.
- Compensation of up to $30 per hour equivalent, varying by contribution level, project scope, complexity, and expertise.
Tech Stack
Categories
Data Engineering
About Mindrift
Mindrift builds an expert-sourcing platform that connects domain specialists to project-based work training and evaluating generative AI models, including supervised fine-tuning, RLHF, evaluation, and red-teaming. It is built and operated by Toloka, part of Nebius Group, and run from Amsterdam, Netherlands. Work is fully remote and freelance, serving global technology companies developing and improving large AI systems.
