Seeking Senior Freelance Web Scraping, Proxy & AI-Directed Data Acquisition Platform Architect/Developer (Kolkata)

Seeking Senior Freelance Web Scraping, Proxy & AI-Directed Data Acquisition Platform Architect/Developer (Kolkata)

14 Aug
|
Outthinc Global Communications
|
Kolkata

14 Aug

Outthinc Global Communications

Kolkata

Build From Scratch | Freelance / Long-Term Collaboration | Profit-Sharing Opportunity for the Right Partner

We are looking for a senior-level expert who can architect and develop a complete web data acquisition platform from scratch for our upcoming AI-Powered Market Intelligence Platform.

This is not a basic web-scraping project. We need someone who can build the underlying infrastructure for AI-directed, targeted and scalable data collection.

Core Architecture The intended workflow is:

Business Objective → LLM/AI Planning → Data Acquisition Platform → Search/Google Advanced Search/Dork Discovery → URL/Source Discovery → Proxy Management → Scrapers/APIs → Targeted Data Extraction → Validation → Deduplication → Knowledge Base → AI/ML

We already have separate LLM/AI experts. The selected developer will work closely with them to convert AI-generated research/data requirements into executable, reliable scraping jobs.

Key Responsibilities The expert must be capable of developing:

- Complete scraping platform from scratch
- Scrapy / Playwright / Selenium / HTTP-based scraping infrastructure
- Proxy management layer with proxy pools, assignment, health monitoring, failover and multiple provider integration
- Scraping orchestrator for job queues, scheduling, concurrency, retries, workers and monitoring
- Google advanced-search / Dork-style discovery and search-result/URL extraction using appropriate and compliant mechanisms
- Source/domain discovery and filtering
- Modular connectors for eBay, Amazon, Shopify, Etsy, Reddit, YouTube review platforms and other sources
- Targeted field-level extraction rather than unnecessary bulk scraping
- Data validation, normalization and deduplication
- Source/URL/query/data lineage
- Structured database and data pipeline




- APIs through which our LLM team can create, monitor, modify and retrieve scraping jobs
- Monitoring, logging and error handling

Critical Requirement: AI-Directed Scraping The system must allow our LLM layer to determine before scraping:

- What to collect
- Where to collect it
- Which queries/keywords to use
- Which sources/URLs are relevant
- Which fields are required
- Filters/timeframes
- Maximum/sufficient data volume
- When to stop

Therefore

LLM Planning → Targeted Collection → Validation → AI Analysis rather than:

Scrape Everything → Store Everything → AI Filters It The objective is to reduce data volume, scraping/proxy costs, storage, processing and LLM token consumption while improving relevance and precision.

LLM Integration

Our LLM experts will handle:

- LLMs and reasoning
- Research intelligence
- RAG
- AI agents
- Market/product/review intelligence
- Recommendations

The selected developer will handle the data acquisition infrastructure and collaborate with the LLM team on:

- AI-to-scraper APIs
- Data schemas
- Collection specifications
- Intelligent stop conditions
- Feedback loops
- Additional targeted data requests

The architecture should support:

AI → Data → AI → Additional Data → AI

Required Experience

Solid practical experience in

- Large-scale web scraping
- Python / Scrapy / Playwright / Selenium
- Proxy infrastructure and proxy management




- Search/query-driven data discovery
- Google advanced-search/Dork-based workflows
- Distributed scraping / job orchestration
- REST APIs
- PostgreSQL / Redis
- Docker / Linux / cloud infrastructure
- Data validation and deduplication
- Scalable data pipelines

Experience with marketplaces, review platforms, search engines, market intelligence or AI/LLM integration is strongly preferred.

Important

We are not looking for a freelancer who can only write individual scrapers.

We need someone capable of:

Architecting → Developing → Integrating → Scaling the complete Search Proxy Scraping Data Acquisition Infrastructure from scratch.

The system must be designed for legitimate and responsible collection of publicly accessible/authorised data and must not rely on bypassing authentication, security controls or other access restrictions.

MVP The MVP will initially be validated using a controlled source/dataset such as eBay, together with search/discovery functionality, and then expanded to additional marketplaces, review sources and other data sources.

Application Requirements

Please provide

1. 2–3 relevant projects involving large-scale scraping/proxy infrastructure.
2. Your proposed technical architecture for the above system.
3. Recommended technology stack.
4. Approach for AI/LLM → scraping integration.
5. Approach for proxy management and cost optimisation.
6. Estimated MVP timeline and development cost.
7. Explanation of how you would scale from 10,000 → 100,000 → 1M records.

Position

Senior Web Scraping, Proxy & AI-Directed Data Acquisition Platform Architect/Developer

Project: Build from Scratch | MVP → Long-Term Platform | Close collaboration with existing LLM/AI team

📌 Seeking Senior Freelance Web Scraping, Proxy & AI-Directed Data Acquisition Platform Architect/Developer (Kolkata)
🏢 Outthinc Global Communications
📍 Kolkata

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: seeking senior freelance web scraping, proxy & ai-directed data acquisition platform architect/developer (kolkata) / kolkata

Subscribe to this job alert:

Get the latest job offers by email for: seeking senior freelance web scraping, proxy & ai-directed data acquisition platform architect/developer (kolkata) / kolkata