SME-Web Scraping and Crawling — Web Data Platform (Hyderabad)

SME-Web Scraping and Crawling — Web Data Platform (Hyderabad)

24 Sep
|
Three Across
|
Hyderabad

24 Sep

Three Across

Hyderabad

: Subject Matter Expert Web Scraping and Crawling Web Data Platform

Role summary: Own and extend a production-bound POC that combines resilient web-crawling engineering with an AWS Bedrock-based LLM verification agent, in a healthcare-data-compliance-sensitive context.

Responsibilities

Replace scraped SERPs with a paid search API behind the existing search_domain() interface.

Move fetch.py to async httpx with per-domain concurrency limits and politeness budgets.

Introduce proxy rotation and egress management; retire the single-IP failure mode.

Make failure loud: typed outcomes, structured logs, metrics, drift and volume alerts.

Reverse-engineer payer XHR/JSON endpoints to replace Playwright recipes wherever possible.

Honor robots.txt Disallow and Crawl-delay; build a per-domain ToS and licensing register.

Containerize and schedule the pipeline; add CI running the offline tests on every change.

Add HAR replay and golden-file parser tests on real payer HTML, plus daily canary crawls.

- Skills

Area

Technologies / Capabilities

Level

Web & protocol fundamentals

HTTP/1.1, HTTP/2, HTTP/3, TLS, JA3/JA4 fingerprinting, cookies/sessions, OAuth2/JWT/SSO, DOM, encodings

[Must]

Legacy stack (real mileage)

urllib, requests, mechanize, cURL/wget, BeautifulSoup, lxml, html5lib, XPath/XSLT, Scrapy (middlewares, pipelines, CrawlSpider), Selenium 3 patterns, Splash/PhantomJS, RSS/Atom/SOAP/XML feeds, ASP.NET __VIEWSTATE postbacks, frame & table-layout scraping, sitemap.xml, WARC/Common Crawl

[Must]

Modern stack

Playwright (contexts, tracing, network interception),



Puppeteer, Selenium 4/CDP, httpx/aiohttp asyncio, curl_cffi TLS impersonation, selectolax, scrapy-playwright, Crawlee/Apify, JSON-LD/microdata harvesting

[Must]

Reverse engineering

Private/undocumented JSON APIs, XHR & fetch interception, GraphQL, mobile API capture (mitmproxy/Charles), bundled-JS deobfuscation, WASM challenge analysis

[Must]

Methodology breadth

API-first vs browser rendering, static vs JS-rendered SPA (__NEXT_DATA__, Nuxt payloads), BFS/DFS frontier management, URL canonicalization & dedup (bloom filters/hashing), full-refresh vs incremental/CDC crawling, distributed queue-based crawling (Redis/Kafka/SQS), authenticated sessions, pagination & infinite scroll, per-domain politeness

[Must]

Non-HTML extraction

PDF (pdfplumber, PyMuPDF in use today), OCR (Tesseract/cloud), Office formats, LLM/vision-assisted extraction and its cost & failure modes

[Must]

Anti-bot & reliability

Cloudflare, Akamai, DataDome, PerimeterX, Imperva, Kasada; browser/canvas/WebGL fingerprinting; residential/datacenter/mobile proxy rotation; CAPTCHA landscape; honeypot detection; backoff with jitter, circuit breakers, idempotent retries

[Must]

Data engineering





Python expert (async, typing, profiling), SQL, pydantic validation (in use); pandera, entity resolution, Postgres/Mongo/S3/Parquet, Airflow/Prefect/Dagster, Docker/K8s, CI/CD to introduce; Node/TypeScript helpful

[Must]

Testing & observability vcrpy/responses, HAR replay, golden-file parser tests, canary crawls, selector-drift alerting, volume-anomaly detection, structured logging & metrics

[Must]

Build vs. buy

Evaluating Zyte, Bright Data, Oxylabs, Apify, ScrapingBee, Firecrawl on cost, coverage, and risk

[Preferred]

Legal & ethical robots.txt and rate-limit adherence, ToS/CFAA awareness, GDPR & PII minimization, licensed-API-first sourcing; designs defensible, low-risk acquisition strategies in partnership with Legal

[Must]

Experience - Required

8 years in data acquisition; 5 years owning a scraping platform end to end.

Proven scale: 1,000 distinct domains or 10M pages/month sustained in production.

Rescued a brittle legacy scraper, with before/after reliability and cost numbers.

Mentored engineers; set crawl standards, review practice, and on-call runbooks.

Degree optional equivalent practical experience is fully accepted.

Nice-to-Have

US payer policy, formulary, or prior-authorization document domain knowledge.

Open-source contributions to Scrapy, Playwright, Crawlee, or a parsing library.

LLM-assisted extraction at controlled cost we run AWS Bedrock in verifier/.

Compliance or legal-review exposure on data acquisition programmes.

📌 SME-Web Scraping and Crawling — Web Data Platform (Hyderabad)
🏢 Three Across
📍 Hyderabad

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: sme-web scraping and crawling — web data platform (hyderabad) / hyderabad

Subscribe to this job alert:

Get the latest job offers by email for: sme-web scraping and crawling — web data platform (hyderabad) / hyderabad