4 days ago
Remote, Japan or Singapore, SingaporeSenior
Responsibilities
- Own end-to-end data extraction workflows across complex websites and deliver accurate, complete, structured datasets.
- Use custom workflows and tools to accelerate data collection, validation, and task execution.
- Extract data from dynamic and interactive websites, including JavaScript-rendered content, while adapting to changing site behavior.
- Apply validation, cross-source consistency checks, formatting controls, and systematic verification to maintain data quality.
- Scale scraping operations through batching or parallelization, monitor failures, and maintain stability against site changes.
Requirements
- At least 5+ years of relevant experience in data engineering, web scraping, automation, or software development.
- Strong Python web-scraping experience with BeautifulSoup, Selenium or similar tools, including dynamic content and APIs via proxies.
- Ability to extract data from complex structures such as hierarchies, archived pages, and inconsistent HTML.
- Experience with data cleaning, normalization, validation, and delivery of structured CSV, JSON, or Google Sheets datasets.
- Experience handling anti-bot mechanisms and dynamic site structures at scale.
- Experience with AWS or equivalent cloud infrastructure and Docker containerization.
- Hands-on experience with LLM frameworks such as LangChain or OpenRouter applied to automation tasks.
- Upper-intermediate English proficiency at B2 level or above is required.
- Strong technical problem-solving, independent troubleshooting, attention to detail, and data-accuracy skills are required; a GitHub link is a plus.
Benefits
- Remote, part-time freelance work.
- Estimated workload of approximately 10–20 hours per week during active project phases, with no guaranteed workload.
- Work on the Tendem project involving specialized data-scraping workflows for real-world applications.
Tech Stack
Categories
Data Engineering
