Senior Python Data Scraping Engineer
Posted 20hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Senior Python engineer building reliable web-scraping workflows for Mindrift’s AI platform. Extracting, validating, and scaling structured datasets from complex dynamic websites.
Responsibilities:
- Own end-to-end data extraction workflows across complex websites, ensuring complete coverage, accuracy, and reliable delivery of structured datasets
- Leverage available tools and custom workflows to accelerate data collection, validation, and task execution while meeting defined requirements
- Ensure reliable extraction from dynamic and interactive web sources, adapting approaches for JavaScript-rendered content and changing site behavior
- Enforce data quality standards through validation checks, cross-source consistency controls, formatting specifications, and systematic verification before delivery
- Scale scraping operations for large datasets using efficient batching or parallelization
- Monitor failures and maintain stability against minor site structure changes
- Collaborate with Tendem Agents handling repetitive tasks within Mindrift’s hybrid AI + human system
- Use tools such as Apify, OpenRouter, and other technologies alongside technical expertise and custom approaches
- Apply web scraping, data extraction, and data processing expertise to deliver accurate, reliable, high-quality results
Requirements:
- At least 5+ years of relevant experience in data engineering, web scraping, automation, or software development (required)
- Bachelor’s or Master’s Degree in Engineering, Applied Mathematics, Computer Science, or related technical fields is a plus
- Strong experience in Python web scraping (BeautifulSoup, Selenium or similar), including dynamic content (JS, AJAX, infinite scroll) and APIs via proxies
- Proven ability to extract data from complex structures, including hierarchies, archived pages, and inconsistent HTML
- Solid background in data cleaning, normalization, and validation, delivering structured datasets in CSV, JSON, or Google Sheets
- Demonstrated experience handling anti-bot mechanisms and dynamic site structures at scale
- Experience with cloud infrastructure (AWS or equivalent) and containerization (Docker) as part of real workflows
- Hands-on experience with LLM frameworks (LangChain, OpenRouter, or similar) applied to automation tasks
- Strong attention to detail and commitment to data accuracy
- Self-directed work ethic with ability to troubleshoot independently
- English proficiency: Upper-intermediate (B2) or above (required)
- A link to GitHub is a plus
Benefits:
- Freelance opportunity
- Part-time remote work
- Estimated 10–20 hours per week during active project phases




