14 Aug
|
Outthinc Global Communications
|
Kolkata
14 Aug
Outthinc Global Communications
Kolkata
Build From Scratch | Freelance / Long-Term Collaboration | Profit-Sharing Opportunity for the Right Partner
We are looking for a senior-level expert who can architect and develop a complete web data acquisition platform from scratch for our upcoming AI-Powered Market Intelligence Platform.
This is not a basic web-scraping project. We need someone who can build the underlying infrastructure for AI-directed, targeted and scalable data collection.
Core Architecture The intended workflow is:
Business Objective → LLM/AI Planning → Data Acquisition Platform → Search/Google Advanced Search/Dork Discovery → URL/Source Discovery → Proxy Management → Scrapers/APIs → Targeted Data Extraction → Validation → Deduplication → Knowledge Base → AI/ML
We already have separate LLM/AI experts. The selected developer will work closely with them to convert AI-generated research/data requirements into executable, reliable scraping jobs.
Key Responsibilities The expert must be capable of developing:
- Complete scraping platform from scratch
- Scrapy / Playwright / Selenium / HTTP-based scraping infrastructure
- Proxy management layer with proxy pools, assignment, health monitoring, failover and multiple provider integration
- Scraping orchestrator for job queues, scheduling, concurrency, retries, workers and monitoring
- Google advanced-search / Dork-style discovery and search-result/URL extraction using appropriate and compliant mechanisms
- Source/domain discovery and filtering
- Modular connectors for eBay, Amazon, Shopify, Etsy, Reddit, YouTube review platforms and other sources
- Targeted field-level extraction rather than unnecessary bulk scraping
- Data validation, normalization and deduplication
- Source/URL/query/data lineage
- Structured database and data pipeline
- APIs through which our LLM team can create, monitor, modify and retrieve scraping jobs
- Monitoring, logging and error handling
Critical Requirement: AI-Directed Scraping The system must allow our LLM layer to determine before scraping:
- What to collect
- Where to collect it
- Which queries/keywords to use
- Which sources/URLs are relevant
- Which fields are required
- Filters/timeframes
- Maximum/sufficient data volume
- When to stop
Therefore
LLM Planning → Targeted Collection → Validation → AI Analysis rather than:
Scrape Everything → Store Everything → AI Filters It The objective is to reduce data volume, scraping/proxy costs, storage, processing and LLM token consumption while improving relevance and precision.
LLM Integration
Our LLM experts will handle:
- LLMs and reasoning
- Research intelligence
- RAG
- AI agents
- Market/product/review intelligence
- Recommendations
The selected developer will handle the data acquisition infrastructure and collaborate with the LLM team on:
- AI-to-scraper APIs
- Data schemas
- Collection specifications
- Intelligent stop conditions
- Feedback loops
- Additional targeted data requests
The architecture should support:
AI → Data → AI → Additional Data → AI
Required Experience
Solid practical experience in
- Large-scale web scraping
- Python / Scrapy / Playwright / Selenium
- Proxy infrastructure and proxy management
- Search/query-driven data discovery
- Google advanced-search/Dork-based workflows
- Distributed scraping / job orchestration
- REST APIs
- PostgreSQL / Redis
- Docker / Linux / cloud infrastructure
- Data validation and deduplication
- Scalable data pipelines
Experience with marketplaces, review platforms, search engines, market intelligence or AI/LLM integration is strongly preferred.
Important
We are not looking for a freelancer who can only write individual scrapers.
We need someone capable of:
Architecting → Developing → Integrating → Scaling the complete Search Proxy Scraping Data Acquisition Infrastructure from scratch.
The system must be designed for legitimate and responsible collection of publicly accessible/authorised data and must not rely on bypassing authentication, security controls or other access restrictions.
MVP The MVP will initially be validated using a controlled source/dataset such as eBay, together with search/discovery functionality, and then expanded to additional marketplaces, review sources and other data sources.
Application Requirements
Please provide
1. 2–3 relevant projects involving large-scale scraping/proxy infrastructure.
2. Your proposed technical architecture for the above system.
3. Recommended technology stack.
4. Approach for AI/LLM → scraping integration.
5. Approach for proxy management and cost optimisation.
6. Estimated MVP timeline and development cost.
7. Explanation of how you would scale from 10,000 → 100,000 → 1M records.
Position
Senior Web Scraping, Proxy & AI-Directed Data Acquisition Platform Architect/Developer
Project: Build from Scratch | MVP → Long-Term Platform | Close collaboration with existing LLM/AI team
📌 Seeking Senior Freelance Web Scraping, Proxy & AI-Directed Data Acquisition Platform Architect/Developer (Kolkata)
🏢 Outthinc Global Communications
📍 Kolkata