Smart Bhujal is a specialist data-engineering and water-intelligence consultancy. We build automated, compliant data pipelines and platforms for research institutes and development organisations. We are currently delivering the backend data-acquisition workstream for a flagship global water-intelligence platform. ←
*client and conference reference removed here*
The role
We are hiring a Senior Crawler & Data-Extraction Engineer to own the core data-acquisition engine of the platform: a compliant discovery crawler and an AI-assisted extraction pipeline that turn fragmented public web sources — government portals, river-basin-authority sites, policy databases and PDFs — into structured, provenance-tracked profiles. This is the most technically demanding role on the team; your work feeds every other module, so it must be reliable, compliant and delivered to schedule.
Location & engagement
Employment type: Contract, fixed-term for the assignment (~4.5 months, extendable)
Location:
Remote (India-based)
Level: Senior · 5–8 years
Reports to: Technical Lead
Start: Immediate
Key responsibilities
• Design and build a discovery and crawling pipeline over authoritative seed lists, designed for global coverage and delivered first for an African pilot.
• Implement a compliance engine — robots.txt and terms-of-use handling, conservative rate limiting, caching, identifiable user-agent, graceful back-off, off-peak scheduling and proxy rotation — so crawling is polite and low-impact.
• Build LLM-assisted extraction from HTML and PDF into structured fields, with multilingual handling, confidence scoring and a low-confidence human-review queue (uncertain data is flagged, never silently written).
• Classify data sources as official vs unofficial and capture full source metadata and provenance for every value.
• Implement staged policy-document fetching with metadata extraction and