{"repo":"adbar/courlan","free":true,"listed":false,"github":"https://github.com/adbar/courlan","clone":"git clone https://github.com/adbar/courlan.git","description":"Clean, filter and sample URLs to optimize data collection – Python & command-line – Deduplication, spam, content and language filters","language":"Python","stars":178,"topics":["url","url-parsing","crawler","tld","uri","url-validation","url-parser","recon","crawling","url-checker"],"license":"Apache-2.0","category":"scrapers-browser-automation","readme_excerpt":"coURLan: Clean, filter, normalize, and sample URLs Why coURLan? \"It is important for the crawler to visit 'important' pages first, so that the fraction of the Web that is visited (and kept up to date) is more meaningful.\" (Cho et al. 1998) \"Given that the bandwidth for conducting crawls is neither infinite nor free, it is becoming essential to crawl the Web in not only a scalable, but efficient way, if some reasonable measure of quality or freshness is to be maintained.\" (Edwards et al. 2001) This library provides an additional \"brain\" for web crawling, scraping and document management. It facilitates web navigation through a set of filters, enhancing the quality of resulting document collections: - Save bandwidth and processing time by steering clear of pages deemed low-value - Identify specific pages based on language or text content - Pinpoint pages relevant for efficient link gathering Additional utilities needed include URL storage, filtering, and deduplication. Features Separate the wheat from the chaff and optimize document discovery and retrieval: - URL handling - Validation - Normalization - Sampling - Heuristics for link filtering - Spam, trackers, and content-types - Locales and internationalization - Web crawling (frontier, scheduling) - Data store specifically designed for URLs - Usable with Python or on the command-line Let the coURLan fish up juicy bits for you! Here is a courlan (source: Limpkin at Harn's Marsh by Russ.jpg), CC BY 2.0). Installation This packa","default_branch":null,"files":null,"tree":[],"storefront":"/r/adbar","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/adbar/courlan/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}