SME-Web Scraping and Crawling — Web Data Platform (Hyderabad)

SME-Web Scraping and Crawling — Web Data Platform (Hyderabad)

24 Sep
|
Three Across
|
Hyderabad

24 Sep

Three Across

Hyderabad

: Subject Matter Expert Web Scraping and Crawling Web Data Platform Role summary: Own and extend a production-bound POC that combines resilient web-crawling engineering with an AWS Bedrock-based LLM verification agent, in a healthcare-data-compliance-sensitive context.

Responsibilities Replace scraped SERPs with a paid search API behind the existing search_domain() interface.

Move fetch.py to async httpx with per-domain concurrency limits and politeness budgets.

Introduce proxy rotation and egress management; retire the single-IP failure mode.

Make failure loud: typed outcomes, structured logs, metrics, drift and volume alerts.

Reverse-engineer payer XHR/JSON endpoints to replace Playwright recipes wherever possible.

Honor robots.txt Disallow and Crawl-delay; build a per-domain ToS and licensing register.

Containerize and schedule the pipeline; add CI running the offline tests on every change.

Add HAR replay and golden-file parser tests on real payer HTML, plus daily canary crawls.

- Skills

Area Technologies / Capabilities Level Web & protocol fundamentals HTTP/1.1, HTTP/2, HTTP/3, TLS, JA3/JA4 fingerprinting, cookies/sessions, OAuth2/JWT/SSO, DOM, encodings [Must]

Legacy stack (real mileage) urllib, requests, mechanize, cURL/wget, BeautifulSoup, lxml, html5lib, XPath/XSLT, Scrapy (middlewares, pipelines, CrawlSpider), Selenium 3 patterns, Splash/PhantomJS, RSS/Atom/SOAP/XML feeds, ASP.NET __VIEWSTATE postbacks, frame & table-layout scraping, sitemap.xml, WARC/Common Crawl

[Must]

Modern stack Playwright (contexts, tracing, network interception), Puppeteer, Selenium 4/CDP, httpx/aiohttp + asyncio, curl_cffi TLS impersonation, selectolax, scrapy-playwright, Crawlee/Apify, JSON-LD/microdata harvesting

[Must]

Reverse engineering Private/undocumented JSON APIs, XHR & fetch interception, GraphQL, mobile API capture (mitmproxy/Charles), bundled-JS deobfuscation, WASM challenge analysis

[Must]

Methodology breadth API-first vs browser rendering, static vs JS-rendered SPA (__NEXT_DATA__, Nuxt payloads), BFS/DFS frontier management,



URL canonicalization & dedup (bloom filters/hashing), full-refresh vs incremental/CDC crawling, distributed queue-based crawling (Redis/Kafka/SQS), authenticated sessions, pagination & infinite scroll, per-domain politeness

[Must]

Non-HTML extraction PDF (pdfplumber, PyMuPDF in use today), OCR (Tesseract/cloud), Office formats, LLM/vision-assisted extraction and its cost & failure modes

[Must]

Anti-bot & reliability Cloudflare, Akamai, DataDome, PerimeterX, Imperva, Kasada; browser/canvas/WebGL fingerprinting; residential/datacenter/mobile proxy rotation; CAPTCHA landscape; honeypot detection; backoff with jitter, circuit breakers, idempotent retries

[Must]

Data engineering Python expert (async, typing, profiling), SQL, pydantic validation (in use); pandera, entity resolution, Postgres/Mongo/S3/Parquet, Airflow/Prefect/Dagster, Docker/K8s, CI/CD to introduce; Node/TypeScript practical

[Must]

Testing & observability vcrpy/responses, HAR replay, golden-file parser tests, canary crawls, selector-drift alerting, volume-anomaly detection, structured logging & metrics

[Must]

Build vs. buy Evaluating Zyte, Bright Data, Oxylabs, Apify, ScrapingBee, Firecrawl on cost, coverage, and risk

[Preferred]

Legal & ethical robots.txt and rate-limit adherence, ToS/CFAA awareness, GDPR & PII minimization, licensed-API-first sourcing; designs defensible, low-risk acquisition strategies in partnership with Legal

[Must]

Experience - Required 8+ years in data acquisition; 5+ years owning a scraping platform end to end.

Proven scale: 1,000+ distinct domains or 10M+ pages/month sustained in production.

Rescued a brittle legacy scraper, with before/after reliability and cost numbers.

Mentored engineers; set crawl standards, review practice, and on-call runbooks.

Degree optional equivalent practical experience is fully accepted.

Nice-to-Have US payer policy, formulary, or prior-authorization document domain knowledge.

Open-source contributions to Scrapy, Playwright, Crawlee, or a parsing library.

LLM-assisted extraction at controlled cost we run AWS Bedrock in verifier/.

Compliance or legal-review exposure on data acquisition programmes.

📌 SME-Web Scraping and Crawling — Web Data Platform (Hyderabad)
🏢 Three Across
📍 Hyderabad

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: sme-web scraping and crawling — web data platform (hyderabad) / hyderabad

Subscribe to this job alert:

Get the latest job offers by email for: sme-web scraping and crawling — web data platform (hyderabad) / hyderabad