7 days ago
Lisbon, PortugalSenior
Responsibilities
- Design and implement advanced web-scraping solutions using Python, Playwright, and AWS services.
- Automate interactions with JavaScript-heavy and authentication-secured websites, including MFA, CAPTCHAs, and session or token-based login flows.
- Architect serverless scraping pipelines using AWS Lambda, Step Functions, S3, CloudWatch, and Secrets Manager.
- Build high-volume data-extraction systems with fault tolerance, retries, and intelligent logging.
- Integrate workflows across multiple portals, APIs, and data sources.
- Contribute to architecture decisions, tooling, and engineering best practices.
Requirements
- 5+ years of Python development experience focused on automation and data extraction.
- Expertise with web scraping tools including Playwright, Selenium, Scrapy, BeautifulSoup, and requests.
- Experience handling multi-step authentication flows, MFA, CAPTCHA solving, and session or cookie management.
- Proficiency deploying and managing workloads in AWS, particularly Lambda, S3, IAM, CloudWatch, and Secrets Manager.
- Experience with asynchronous programming, headless browsers, and JavaScript-rendered content.
- Understanding of HTTP, HTTPS, cookies, headers, network-call analysis, and authentication mechanisms.
- Experience working with APIs, JSON/XML, and data transformation.
- Bachelor’s or Master’s degree in Computer Science, Software Engineering, or a related field, or equivalent professional experience.
- Preferred experience includes CI/CD pipelines, Docker, CloudFormation, Terraform, Apache Airflow or AWS Step Functions, secure enterprise portals, data engineering, ETL workflows, and Python testing frameworks.
Benefits
- Comprehensive benefits package with healthcare and 100% employer-paid health and dental insurance.
- Competitive salary, annual performance bonus, and equity for full-time employees.
- Generous paid time off.
- Hybrid work arrangement with four days in the office and remote work flexibility on Fridays.
