19 Sep
|
AIMLEAP
|
Bengaluru
Job DescriptionPython & JavaScript Developer – AI & Web Scraping
NExperience: 4-7 Years
NLocation: Remote
NMode of Engagement: Full-time
NNo of Positions: 3
NEducational Qualifications: Bachelor's degree in computer science, Information Technology
NIndustry: IT / Software Development
NNotice Period: Immediate Joiners Preferred
NAbout the Role
nWeare looking for a hands-on Web Scraping / Crawling Engineer with 3–5 years of experience in web scraping, browser automation, and scalable data extraction.
NThe ideal candidate should have strong experience working with dynamic and JavaScript-heavy websites, building reliable crawling workflows, handling crawling failures, and working with distributed processing systems.
NThis is a highly technical, hands-on role. You will be expected to design, develop, debug, optimize, and maintain web crawling and data extraction systems.
NResponsibilities
n
N
- Design, develop, and maintain scalable web crawling and scraping systems for dynamic and JavaScript-heavy websites.
N
- Develop browser automation workflows using Playwright, Selenium, Puppeteer, or similar frameworks.
N
- Investigate and resolvecrawling issues such as 403/429 responses, redirects, timeouts, rendering failures, session issues, and anti-bot challenges.
N
- Develop robust retry, fallback, and failure-handling mechanisms to improve crawler reliability.
N
- Work with cookies, sessions, browser contexts, headers, proxies, and related crawling mechanisms to maintain state and improve crawling reliability.
N
- Build and optimize concurrent and distributed scraping workflows using asynchronous processing, queues, and worker-based architectures.
N
- Design data pipelines covering URL processing, crawling, extraction, validation, transformation, and storage.
N
- Optimize crawler performance, including concurrency, browser lifecycle, resource utilization, timeouts, and request handling.
N
- Implement monitoring and observability for crawl success rates, failure types, latency, retries, and worker performance.
N
- Debug complex crawling problems and identify root causes rather than relying only on one-off fixes.
N
- Develop reusable crawling components and frameworks that can support multiple websites and use cases.
N
- Collaborate with data engineering, AI, backend, and product teams to deliver reliable and structured web data.
N
- Evaluate and adopt new web crawling, browser automation, and data extraction technologies where appropriate.
N
nRequired Skills
N
n
- 3–5 years of hands-on experience in web scraping, web crawling, browser automation, or web data engineering.
N
- Strong programming experience in Python. JavaScript/Node.Js is a plus.
N
- Strong hands-on experience with Scrapy, Playwright, Selenium, Puppeteer, or similar browser automation frameworks.
N
- Strong understanding ofHTTP, HTML, DOM, JavaScript rendering, redirects, cookies, sessions, headers, and browser contexts.
N
- Experience with browserfingerprinting, WAFs, and modern anti-bot mechanisms.
N
- Experience working withdynamic, JavaScript-heavy, and asynchronous websites.
N
- Valuable understanding of asynchronous programming, concurrency, and parallel processing.
N
- Experience with job queues, distributed workers, or message-based processing systems such as Redis, RabbitMQ, Kafka, Celery, or similar technologies.
N
- Understanding of retry mechanisms, error handling, idempotency, rate limiting, and backpressure in distributed systems.
N
- Hands-on experience with proxy management, session handling, and anti-bot challenges.
N
- Strong debugging and analytical skills, with the ability to investigate and resolve complex crawling failures.
N
- Experience working withSQL databases such as PostgreSQL or MySQL.
N
- Familiarity with Dockerand cloud platforms such as AWS, GCP, or Azure is a plus.
N
- Strong hands-on engineering mindset rather than purely project or delivery management experience.
N
- Ability to independently debug, investigate, and solve difficult crawling problems.
N
- Strong understanding ofhow browsers, HTTP requests, sessions, and websites interact.
N
- Ability to design systems that are reliable, scalable, and fault tolerant.
N
- Good understanding of how to distribute and coordinate large-scale scraping workloads.
N
- Ability to identify theroot cause of crawling failures and develop generalized, reusable solutions.
N
- Willingness to work across scraping, backend services, distributed processing, data pipelines, and infrastructure when required.
N
- Strong problem-solving skills and curiosity to understand how websites behave rather than relying solely on existing scraping tools.
N
nEducation
n
N
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field is preferred.
N
- Equivalent practical experience in software engineering or web scraping will also be considered.
N
📌 Hiring: Web Scraping Engineer-Python (Remote) (Bengaluru)
🏢 AIMLEAP
📍 Bengaluru