24 Sep
|
Three Across
|
Hyderabad
24 Sep
Three Across
Hyderabad
: Subject Matter Expert Web Scraping and Crawling Web Data Platform Role summary: Own and extend a production-bound POC that combines resilient web-crawling engineering with an AWS Bedrock-based LLM verification agent, in a healthcare-data-compliance-sensitive context.
Responsibilities Replace scraped SERPs with a paid search API behind the existing search_domain() interface.
Move fetch.py to async httpx with per-domain concurrency limits and politeness budgets.
Introduce proxy rotation and egress management; retire the single-IP failure mode.
Make failure loud: typed outcomes, structured logs, metrics, drift and volume alerts.
Reverse-engineer payer XHR/JSON endpoints to replace Playwright recipes wherever possible.
Honor robots.txt Disallow and Crawl-delay; build a per-domain ToS and licensing register.
Containerize and schedule the pipeline; add CI running the offline tests on every change.
Add HAR replay and golden-file parser tests on real payer HTML, plus daily canary crawls.
- Skills
Area Technologies / Capabilities Level Web & protocol fundamentals HTTP/1.1, HTTP/2, HTTP/3, TLS, JA3/JA4 fingerprinting, cookies/sessions, OAuth2/JWT/SSO, DOM, encodings [Must]
Legacy stack (real mileage) urllib, requests, mechanize, cURL/wget, BeautifulSoup, lxml, html5lib, XPath/XSLT, Scrapy (middlewares, pipelines, CrawlSpider), Selenium 3 patterns, Splash/PhantomJS, RSS/Atom/SOAP/XML feeds, ASP.NET __VIEWSTATE postbacks, frame & table-layout scraping, sitemap.xml, WARC/Common Crawl
[Must]
Modern stack Playwright (contexts, tracing, network interception), Puppeteer, Selenium 4/CDP, httpx/aiohttp + asyncio, curl_cffi TLS impersonation, selectolax, scrapy-playwright, Crawlee/Apify, JSON-LD/microdata harvesting
[Must]
Reverse engineering Private/undocumented JSON APIs, XHR & fetch interception, GraphQL, mobile API capture (mitmproxy/Charles), bundled-JS deobfuscation, WASM challenge analysis
[Must]
Methodology breadth API-first vs browser rendering, static vs JS-rendered SPA (__NEXT_DATA__, Nuxt payloads), BFS/DFS frontier management,
URL canonicalization & dedup (bloom filters/hashing), full-refresh vs incremental/CDC crawling, distributed queue-based crawling (Redis/Kafka/SQS), authenticated sessions, pagination & infinite scroll, per-domain politeness
[Must]
Non-HTML extraction PDF (pdfplumber, PyMuPDF in use today), OCR (Tesseract/cloud), Office formats, LLM/vision-assisted extraction and its cost & failure modes
[Must]
Anti-bot & reliability Cloudflare, Akamai, DataDome, PerimeterX, Imperva, Kasada; browser/canvas/WebGL fingerprinting; residential/datacenter/mobile proxy rotation; CAPTCHA landscape; honeypot detection; backoff with jitter, circuit breakers, idempotent retries
[Must]
Data engineering Python expert (async, typing, profiling), SQL, pydantic validation (in use); pandera, entity resolution, Postgres/Mongo/S3/Parquet, Airflow/Prefect/Dagster, Docker/K8s, CI/CD to introduce; Node/TypeScript practical
[Must]
Testing & observability vcrpy/responses, HAR replay, golden-file parser tests, canary crawls, selector-drift alerting, volume-anomaly detection, structured logging & metrics
[Must]
Build vs. buy Evaluating Zyte, Bright Data, Oxylabs, Apify, ScrapingBee, Firecrawl on cost, coverage, and risk
[Preferred]
Legal & ethical robots.txt and rate-limit adherence, ToS/CFAA awareness, GDPR & PII minimization, licensed-API-first sourcing; designs defensible, low-risk acquisition strategies in partnership with Legal
[Must]
Experience - Required 8+ years in data acquisition; 5+ years owning a scraping platform end to end.
Proven scale: 1,000+ distinct domains or 10M+ pages/month sustained in production.
Rescued a brittle legacy scraper, with before/after reliability and cost numbers.
Mentored engineers; set crawl standards, review practice, and on-call runbooks.
Degree optional equivalent practical experience is fully accepted.
Nice-to-Have US payer policy, formulary, or prior-authorization document domain knowledge.
Open-source contributions to Scrapy, Playwright, Crawlee, or a parsing library.
LLM-assisted extraction at controlled cost we run AWS Bedrock in verifier/.
Compliance or legal-review exposure on data acquisition programmes.
📌 SME-Web Scraping and Crawling — Web Data Platform (Hyderabad)
🏢 Three Across
📍 Hyderabad