11 Sep
|
Krawlnet Technologies
|
Pune
11 Sep
Krawlnet Technologies
Pune
Role Overview
We are looking for a Web Scraping Engineer to design, build, and operate high-reliability web data acquisition systems in dynamic web environments. The role is hands-on and technical, focused on solving real-world scraping challenges such as anti-bot defenses, session management, browser automation, and large-scale parallel crawling.
Key Responsibilities
Core Web Scraping & Data Acquisition Design and implement high-parallel crawlers for complex, dynamic websites
Handle JavaScript-heavy sites using Python / Playwright / Puppeteer / CDP
Implement robust session management, cookies, headers, and identity rotation
Build resilient pipelines for structured & unstructured data extraction
Ensure data accuracy, freshness, and consistency at scale
Anti-Bot & Protection Handling Analyze and mitigate anti-bot protections, including:
Cloudflare
Akamai
Qrator
ServicePipe
Variti (and similar)
Work with browser-level and network-level fingerprinting:
TLS / JA3
CDP fingerprints
Header and behavior-based detection
Integrate and optimize CAPTCHA solving strategies (where legally permitted)
Required Skills & Experience
Must-Have 1–3+ years of hands-on web scraping / crawling experience
Strong proficiency in Python (Scrapy, asyncio, aiohttp, Playwright)
Deep understanding of:
HTTP, cookies, headers, redirects
Browser automation & JS execution
Experience scraping highly protected websites
Solid debugging skills in hostile scraping environments
Nice-to-Have (Not Mandatory) Familiarity with:
Kubernetes / Helm (user-level, not platform owner)
CI/CD pipelines
Distributed task queues (Redis, RabbitMQ, Kafka)
Experience mentoring junior engineers
Role Details
Role: Data Engineer
Industry Type: IT Services & Consulting
Department: Engineering – Software & QA
Employment Type: Full Time, Permanent
Role Category: Software Development
Education UG: B.Tech / B.E. in Any Specialization
Key Skills
Scrapy
Web Crawling
Python
Data Scraping
Data Extraction
About Company
KrawlNet Technologies enables advertisers and publishers to run affiliate programs efficiently and at scale. We aggregate product data from a wide range of retailers, allowing publishers and analytics partners to seamlessly access high-quality, normalised data to drive growth and performance.
At the core of our platform is web-scale crawling and data extraction, built to handle complex, agile web environments. Our mission is to solve critical industry challenges by delivering reliable services for data acquisition, cleansing, normalisation, and enrichment, transforming raw web content into actionable business intelligence.
Company Info Address: Suratwala Mark Plazzo, Hinjawadi, Pune, Maharashtra 411057
📌 Web Scraping Engineer in Pune
🏢 Krawlnet Technologies
📍 Pune