16 Aug
|
Merit Data and Technology
|
India
16 Aug
Merit Data and Technology
India
Experience : 12 - 18 Years
Location : PAN INDIA
Work Mode : Permanent Remote
Job Type : Full Time Employment
Notice : Immediate Joiners
:
Primary (70%) : Web scraping architecture anti-bot, distributed crawling, proxy/CAPTCHA strategy, compliance.
Secondary (30%) : Data engineering pipelines, lakehouse, orchestration, dbt.
Mandatory ownership of the technical solution and effort estimation for every current scraping proposal/RFP.
Senior profile : 9+ years overall, 5+ years in scraping, 4+ years in data engineering.
Technical Lead / Architect Web Scraping & Data Engineering :
Role Summary :
We are seeking a Technical lead for the design and delivery of large-scale web scraping and data extraction solutions, with strong supporting expertise in data engineering, the technical authority on all scraping initiatives from pre-sales solutioning and effort estimation through to architecture, build, and stabilization.
The primary mandate is web scraping : designing resilient crawlers, anti-bot strategies, distributed extraction systems, and compliance frameworks. The secondary mandate is data engineering : ensuring extracted data flows reliably into well-modelled, query-ready storage layers and downstream analytical or operational systems.
Mandatory Involvement Areas :
- Author or co-author the technical solution section of every new RFP / RFI / proposal document for scraping projects.
- Lead technical discovery calls with prospective clients to understand target sources, data SLAs, volume, and compliance constraints.
- Conduct target-site feasibility assessments anti-bot complexity, dynamic content, login walls, geo-restrictions, rate limits and document findings before commercials are committed.
- Own end-to-end effort estimation for scraping engagements : crawler build effort, infrastructure sizing, proxy and CAPTCHA cost projections, maintenance overhead, and contingency.
- Produce solution architecture diagrams, tech stack recommendations, and assumption logs as part of every proposal.
- Define and document SLAs, KPIs, and acceptance criteria proposed to the client.
- Participate in client orals, technical defence sessions, and commercial negotiations as the technical SPOC.
- Maintain an internal estimation knowledge base reusable estimation templates, complexity matrices, target-site classification, and historical actuals and continuously refine it after every project closure.
- Sign off on the technical feasibility and risk profile of every proposal before submission. No scraping proposal goes out without architect approval.
Core Responsibilities :
- Design end-to-end scraping solutions covering crawl orchestration, extraction, parsing, storage, and downstream data consumption.
- Define reference architectures for high-volume, high-velocity, and high-variety scraping use cases.
- Architect resilient systems handling JavaScript-heavy sites, CAPTCHAs, rate limiting, IP blocking, and frequent DOM changes.
- Evaluate and select frameworks, proxy networks, and anti-bot bypass strategies aligned with cost, performance, and compliance.
- Design data quality, deduplication, validation, and schema-evolution strategies for scraped data.
Data Engineering Architecture (Secondary) :
- Design downstream data pipelines that ingest, transform, and serve scraped data to analytics, ML, or operational consumers.
- Architect lakehouse / warehouse layers, define data modelling standards (dimensional, Data Vault, or hybrid), and govern schema evolution.
- Define ELT/ETL patterns, orchestration strategy, and SLA-backed data freshness commitments.
- Establish data quality, observability, lineage, and cataloguing practices across the platform.
- Recommend storage formats, partitioning, and indexing strategies for cost and query performance.
Delivery Leadership :
- Translate business requirements into technical specifications, sprint plans, and implementation roadmaps.
- Produce HLD, LLD, and Architecture Decision Records (ADRs) for each engagement.
- Provide hands-on guidance, perform code reviews, and mentor scraping and data engineers.
- Drive proof-of-concept (PoC) builds for complex or high-risk targets before full-scale rollout.
Compliance, Risk & Governance :
- Ensure all scraping work adheres to applicable laws and platform terms (GDPR, CCPA, DPDP, robots.txt, copyright).
- Define and enforce ethical scraping practices, request throttling, and PII handling guidelines.
- Conduct risk assessments for each target source and recommend mitigations.
Performance, Cost & Operations :
- Establish SLAs for crawl freshness, completeness, and accuracy.
- Design monitoring, alerting, and self-healing mechanisms for scraping pipelines.
- Optimise infrastructure cost (compute, proxies, storage) without compromising delivery KPIs.
Required Technical Skills Primary (Web Scraping) :
These are non-negotiable. The candidate must demonstrate deep, hands-on expertise in each of the following :
Scraping Frameworks & Tooling :
- Expert-level Python : Scrapy, BeautifulSoup, lxml, Requests, httpx, parsel.
- Headless browser automation : Playwright, Puppeteer, Selenium, Pyppeteer.
- Node.js scraping stack (where applicable) : Puppeteer, Cheerio, Crawlee.
Anti-Bot & Evasion Strategy :
- Proven experience bypassing Cloudflare, Akamai Bot Manager, DataDome, PerimeterX, Imperva, Kasada.
- CAPTCHA handling : reCAPTCHA v2/v3, hCaptcha, FunCaptcha, image / audio solvers; integration with 2Captcha, Anti-Captcha, CapSolver.
- TLS / JA3 / JA4 fingerprinting awareness; HTTP/2 fingerprint evasion; browser fingerprint spoofing.
- Stealth plugins, user-agent rotation, header normalisation, cookie / session management at scale.
Proxy & Network Infrastructure :
- Hands-on experience with rotating, residential, mobile, ISP, and datacenter proxies.
- Integration with providers : Bright Data, Oxylabs, Smartproxy, NetNut, IPRoyal, SOAX.
- Proxy pool design, health-checking, geo-targeting, sticky sessions, and cost optimisation.
Parsing & Extraction :
- XPath, CSS selectors, regex, JSON-LD, microdata, RDFa.
- Reverse-engineering of internal / mobile APIs, GraphQL endpoints, and XHR traffic.
- ML/LLM-assisted extraction for unstructured layouts (nice-to-have, increasingly expected).
Distributed Crawling & Orchestration :
- Scrapy-Redis, Scrapy Cluster, Frontera, Crawlee, or equivalent distributed crawling frameworks.
- Job scheduling and orchestration with Apache Airflow, Prefect, Dagster, or Celery.
- Queue-based architectures using Kafka, RabbitMQ, AWS SQS, GCP Pub/Sub.
Required Technical Skills Secondary (Data Engineering) :
Data Pipelines & Orchestration :
- ETL / ELT design patterns, idempotent pipelines, CDC (Change Data Capture), incremental loads.
- Apache Airflow, Prefect, Dagster, AWS Glue, Azure Data Factory, GCP Dataflow.
- Stream processing : Kafka Streams, Apache Flink, Spark Structured Streaming.
- Batch processing : Apache Spark (PySpark), Databricks, EMR, Dataproc.
Data Modelling & Storage :
- Dimensional modelling (Kimball), Data Vault 2.0, normalised vs. denormalised trade-offs.
- Data warehouses : Snowflake, BigQuery, Redshift, Synapse, Databricks SQL Warehouse.
- Data lakes / lakehouses : Delta Lake, Apache Iceberg, Apache Hudi on S3 / GCS / ADLS.
- OLTP databases : PostgreSQL, MySQL; NoSQL : MongoDB, DynamoDB, Cassandra, Elasticsearch, Redis.
- File formats : Parquet, Avro, ORC, JSON, CSV; partitioning, bucketing, compaction strategies.
Transformation & Quality :
- dbt (data build tool) for transformation, testing, and documentation.
- Data quality frameworks : Great Expectations, Soda, Deequ, custom validators.
- Data lineage and cataloguing : DataHub, OpenMetadata, Amundsen, Atlan, Collibra.
Cloud, DevOps & Observability :
- Strong on at least one of AWS, GCP, Azure (compute, storage, IAM, networking, serverless).
- Containerisation and orchestration : Docker, Kubernetes (EKS / GKE / AKS), ECS.
- Infrastructure as Code : Terraform, Pulumi, CloudFormation.
- CI/CD : GitHub Actions, GitLab CI, Jenkins, Argo CD.
- Observability : Prometheus, Grafana, ELK / OpenSearch, Datadog, Sentry, OpenTelemetry.
Experience & Qualifications :
- Bachelor's or Master's degree in Computer Science, Engineering, or related field.
- 10+ years of overall software engineering experience.
- 5+ years dedicated to large-scale web scraping / data extraction (primary).
- 3+ years of hands-on data engineering experience covering pipelines, warehouses / lakehouses, and orchestration (secondary).
- Proven track record of architecting scraping platforms processing millions of pages per day across diverse target sites.
- Prior experience as Solution Architect, Tech Lead, or Principal Engineer leading teams of 5+ engineers.
- Demonstrated experience in pre-sales / proposal authoring / RFP responses for scraping or data engineering engagements.
- Client-facing consulting experience strongly preferred.
Soft Skills :
- Excellent written and verbal communication; able to defend technical proposals to CXO-level audiences.
- Strong commercial acumen understands the cost / risk / quality trade-offs in estimation.
- Analytical, structured problem-solving mindset.
- Ownership-driven, comfortable being the single point of technical accountability.
- High standards for documentation, knowledge transfer, and reusability.
Expected Deliverables (RFP Scope) :
- Technical solution sections of all proposal documents.
- Target-site feasibility and risk assessment reports.
- Effort estimation models, assumption logs, and complexity matrices.
- Solution architecture diagrams and tech stack recommendations for proposals.
Delivery Phase :
- High-Level Design (HLD) and Low-Level Design (LLD) documents per initiative.
- Architecture Decision Records (ADRs).
- Reference implementations / PoCs for complex extraction scenarios.
- Code review records and engineering standards documentation.
- Operational runbooks, monitoring playbooks, and incident response procedures.
- Compliance and data governance documentation.
- Knowledge transfer sessions and final handover documentation.
Nice-to-Have :
- ML / NLP-based extraction, entity resolution, or LLM-assisted parsing experience.
- Exposure to GraphQL, gRPC, mobile API reverse-engineering, mitmproxy / Charles workflows.
- Domain experience : e-commerce price intelligence, market research, financial data, real estate, travel aggregation, hospitality.
- Open-source contributions to scraping or data engineering ecosystems.
- Familiarity with cross-jurisdictional legal frameworks for automated data collection.
- Experience with reverse ETL tools (Hightouch, Census) and feature stores.
📌 Merit Group - Technical Lead - Web Scraping & Data Engineering (India)
🏢 Merit Data and Technology
📍 India