: Subject Matter Expert Web Scraping and Crawling Web Data Platform
Role summary: Own and extend a production-bound POC that combines resilient web-crawling engineering with an AWS Bedrock-based LLM verification agent, in a healthcare-data-compliance-sensitive context.
Responsibilities
- Replace scraped SERPs with a paid search API behind the existing search_domain() interface.
- Move fetch.py to async httpx with per-domain concurrency limits and politeness budgets.
- Introduce proxy rotation and egress management; retire the single-IP failure mode.
- Make failure loud: typed outcomes, structured logs, metrics, drift and volume alerts.
- Reverse-engineer payer XHR/JSON endpoints to replace Playwright recipes wherever possible.
- Honor robots.txt Disallow and Crawl-delay; build a per-domain ToS and licensing register.
- Containerize and schedule the pipeline; add CI running the offline tests on every change.
- Add HAR replay and golden-file parser tests on real payer HTML,
plus daily canary crawls.
Experience - Required
- 8+ years in data acquisition; 5+ years owning a scraping platform end to end.
- Proven scale: 1,000+ distinct domains or 10M+ pages/month sustained in production.
- Rescued a brittle legacy scraper, with before/after reliability and cost numbers.
- Mentored engineers; set crawl standards, review practice, and on-call runbooks.
- Degree optional equivalent practical experience is fully accepted.
Nice-to-Have
- US payer policy, formulary, or prior-authorization document domain knowledge.
- Open-source contributions to Scrapy, Playwright, Crawlee, or a parsing library.
- LLM-assisted extraction at controlled cost we run AWS Bedrock in verifier/.
- Compliance or legal-review exposure on data acquisition programmes.
📌 Consultant - Web Scraping (Hyderabad)
🏢 CHRYSELYS
📍 Hyderabad
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.