—Capture of massive data on the web and mobile terminals, and the design of architectures such as extraction, deduplication, classification, clustering, and filtering;
—Design and development of distributed web crawlers, able to independently solve various problems encountered in the actual development process;
—Research and development of web page information extraction technology algorithms to improve the efficiency and quality of data capture;
—Analysis and warehousing of crawled data, monitoring of the crawler system and abnormal alarms;
—Designing and developing data collection strategies and anti-shielding rules to improve the efficiency and quality of data collection;
—Design and development of core algorithms according to the system data processing flow and business function requirements;
—Own the development of these tools, services, and workflows to enhance data management, crawl/scrape analysis, reports, and workflows
—Control the testing of the data and the scraping to guarantee compliance, quality, and accuracy
—Monitor the procedure to detect and address any problems with breaks and scale scrapes as necessary
—Create systems for handling large and unstructured data while developing a regulatory update tool for legal clients
—Create a tool that gathers information on regulatory updates for legal clients by using scraping bots on websites, especially regulatory websites
Desired Profile / Criteria / Skills :
Qualifications
—Proficient in Python language with JavaScript/React JS/Next Js, familiar with one or more of the commonly used crawler frameworks, such as Scrapy framework or other Web scraping frameworks, with independent development experience
—Have 2+ years of experience with JavaScript and 1.5+ years of experience working with WebScrape, Crawlers, and Data Extraction.
—Familiar with vertical search crawlers and distributed web crawlers, deeply understanding the principles of web crawlers, having rich experienc