{"repo":"fcibecchini/smart-crawler","free":true,"listed":false,"github":"https://github.com/fcibecchini/smart-crawler","clone":"git clone https://github.com/fcibecchini/smart-crawler.git","description":"A smart distributed crawler that infers navigation models of structured websites, used to cluster pages based on their structure and extract data from them.","language":"Java","stars":10,"topics":["crawler","scraper","akka","scraping-websites"],"license":null,"category":"scrapers-browser-automation","readme_excerpt":"Smart crawler of structured websites A smart distributed crawler that infers navigation models of structured websites, used to efficiently crawl their pages and extract structured data from them. The crawling process is divided in 2 phases: 1. Given a list of entrypoints (website homepage URLs), the navigation model of each website is automatically inferred by exploring a limited yet rapresentative sample of theirs HTML pages. The generated models divide each website in classes of similarly structured pages, called Page Classes, and describe the properties of the links between different Page Classes. 2. The so generated models are then used to perform an extensive crawling of the websites, so that each URL can be associated with the Page Class it belongs to: the output consists of webpages clustered according to their HTML structure. Use cases Navigation models and clustered pages can be both useful for different use cases. There you can find some examples: The pages crawled in phase 1., used to infer the model, can be used as samples to generate a wrapper to extract structured data from Page Classes of interest. Inferred models can be used in phase 2. to explore only a portion of a website containing data of interest, rather than all the URLs, by following only the links that will take the crawler to the pages of a Page Class important for the domain of interest (for instance, you want to crawl an e-commerce website but you are interested in all and only the pages showing a ","default_branch":null,"files":null,"tree":[],"storefront":"/r/fcibecchini","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/fcibecchini/smart-crawler/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}