3 months ago
Hyderābād, IndiaMid Level
Responsibilities
- Maintain and enhance scrapers across marketplaces, websites, webshops, domain and WHOIS sources, and social media platforms.
- Build new scrapers from initial site analysis through production deployment.
- Improve scraper reliability against anti-bot defenses using rotating proxies, browser fingerprinting, request pacing, and header and cookie management.
- Monitor scraper health, diagnose failures, and resolve breakages caused by website changes.
- Structure and store collected data so downstream pipelines receive clean, complete, and organized data.
- Run scrapers at scale using scheduling, orchestration, retry, and queue-management mechanisms.
- Package and deploy scrapers in containerized, version-controlled environments.
- Write maintainable code, participate in code reviews, and document scrapers for team use.
Requirements
- At least 3 years of relevant experience in web scraping or a similar domain.
- Strong Python experience, including practical Playwright development.
- Experience with Selenium and scraping at scale against sites with active anti-scraping mechanisms.
- Experience with CAPTCHA-solving approaches, Google Cloud, virtual machines, and marketplace or social media scraping.
- Experience reverse-engineering mobile or web APIs.
- Familiarity with browser fingerprinting libraries such as undetected-chromedriver and playwright-stealth.
- Solid understanding of HTTP, HTML, DOM, browser developer tools, XHR, fetch, and dynamic rendering.
- Experience managing proxies, request headers, cookies, and scraper performance against anti-bot defenses.
- Proficiency with Git and containerized environments such as Docker.
- Exposure to Java or TypeScript is beneficial.
- Fluent written and verbal English.
Categories
BackendData Engineering
