07 Aug
|
Verifitech
|
Chennai
07 Aug
Verifitech
Chennai
Job Title: Senior Data Engineer (Scraping)
Employment Type: Full-Time
Job Summary We are looking for an experienced Senior Data Engineer (Scraping) to design, develop, and maintain scalable web scraping and data harvesting pipelines. The ideal candidate should have solid expertise in Python, web scraping frameworks, ETL processes, distributed data processing, and workflow orchestration. This role involves working closely with the Data Harvest team to collect, process, and deliver high-quality data from multiple web sources and APIs.
Key Responsibilities
- Design, develop, and maintain scalable web scraping and data harvesting pipelines.
- Build and maintain web scrapers using Python frameworks such as Scrapy, BeautifulSoup, Selenium, and Playwright.
- Handle agile JavaScript-rendered websites and overcome anti-bot mechanisms including proxy rotation, IP rotation, user-agent rotation, rate limiting, and CAPTCHA handling.
- Develop ETL workflows for data extraction, parsing, transformation, and cleaning.
- Process and transform large-scale datasets using PySpark and distributed computing.
- Design and schedule workflows using Apache Airflow (Dagster, Prefect, or Luigi experience is an added advantage).
- Store and manage data in SQL and NoSQL databases.
- Work with data formats such as CSV, JSON, XML, and Parquet.
- Implement monitoring, logging, retry mechanisms, and error handling to ensure pipeline reliability.
- Ensure compliance with website policies, robots.txt, and data privacy regulations such as GDPR.
- Collaborate with technical teams and data consumers to maintain data quality and timely delivery.
Requirements
Key Skills
Strong programming experience in Python. Knowledge of Node.js or JavaScript is an added advantage. Hands-on experience with: Scrapy BeautifulSoup Selenium Playwright lxml requests/httpx Puppeteer (preferred) Strong understanding of: HTML CSS DOM XPath CSS Selectors HTTP Protocol Experience with REST APIs and GraphQL APIs. Experience parsing JSON, XML, and HTML data. Strong ETL development experience. Experience with PySpark for distributed data processing. Hands-on experience with Apache Airflow. Good knowledge of SQL and NoSQL databases including PostgreSQL, MySQL, and MongoDB. Experience with asynchronous programming, concurrency, and distributed scraping. Knowledge of Git, Docker, and cloud platforms such as AWS, Azure, or GCP. Experience implementing monitoring and alerting for production pipelines. Understanding of legal and compliance requirements related to web scraping.
Preferred Skills
Experience with Dagster, Prefect, or Luigi. Knowledge of serverless and cloud-native architectures. Experience with large-scale data engineering projects. Strong debugging, troubleshooting, and analytical skills.
Educational Qualification Bachelors or Masters degree in Computer Science, Engineering, Information Technology, or a related field. Equivalent practical experience will also be considered.
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Senior Data Engineer (Scraping) (Chennai)
🏢 Verifitech
📍 Chennai