08 Sep
|
CHRYSELYS
|
Hyderabad
08 Sep
CHRYSELYS
Hyderabad
Job Description: Subject Matter Expert Web Scraping and Crawling Web Data Platform
Role summary: Own and extend a production-bound POC that combines resilient web-crawling engineering with an AWS Bedrock-based LLM verification agent, in a healthcare-data-compliance-sensitive context.
Responsibilities
• Replace scraped SERPs with a paid search API behind the existing search_domain() interface.
• Move fetch.py to async httpx with per-domain concurrency limits and politeness budgets.
• Introduce proxy rotation and egress management; retire the single-IP failure mode.
• Make failure loud: typed outcomes, structured logs, metrics, drift and volume alerts.
• Reverse-engineer payer XHR/JSON endpoints to replace Playwright recipes wherever possible.
• Honor robots.txt Disallow and Crawl-delay; build a per-domain ToS and licensing register.
• Containerize and schedule the pipeline; add CI running the offline tests on every change.
• Add HAR replay and golden-file parser tests on real payer HTML, plus daily canary crawls.
Experience - Required
• 8+ years in data acquisition; 5+ years owning a scraping platform end to end.
• Proven scale: 1,000+ distinct domains or 10M+ pages/month sustained in production.
• Rescued a brittle legacy scraper, with before/after reliability and cost numbers.
• Mentored engineers; set crawl standards, review practice, and on-call runbooks.
• Degree optional equivalent practical experience is fully accepted.
Nice-to-Have
• US payer policy, formulary, or prior-authorization document domain knowledge.
• Open-source contributions to Scrapy, Playwright, Crawlee, or a parsing library.
• LLM-assisted extraction at controlled cost we run AWS Bedrock in verifier/.
• Compliance or legal-review exposure on data acquisition programmes.
📌 Consultant - Web Scraping (Hyderabad)
🏢 CHRYSELYS
📍 Hyderabad