{"repo":"ArchiveTeam/grab-site","free":true,"listed":false,"github":"https://github.com/ArchiveTeam/grab-site","clone":"git clone https://github.com/ArchiveTeam/grab-site.git","description":"The archivist's web crawler: WARC output, dashboard for all crawls, dynamic ignore patterns","language":"Python","stars":1605,"topics":["archiving","crawl","spider","crawler","warc"],"license":null,"category":"scrapers-browser-automation","readme_excerpt":"grab-site ========= [![Build status][travis-image]][travis-url] grab-site is an easy preconfigured web crawler designed for backing up websites. Give grab-site a URL and it will recursively crawl the site and write WARC files. Internally, grab-site uses a fork of wpull for crawling. grab-site gives you a dashboard with all of your crawls, showing which URLs are being grabbed, how many URLs are left in the queue, and more. the ability to add ignore patterns when the crawl is already running. This allows you to skip the crawling of junk URLs that would otherwise prevent your crawl from ever finishing. See below. an extensively tested default ignore set (global) as well as additional (optional) ignore sets for forums, reddit, etc. duplicate page detection: links are not followed on pages whose content duplicates an already-seen page. The URL queue is kept on disk instead of in memory. If you're really lucky, grab-site will manage to crawl a site with 10M pages. Note: if you have any problems whatsoever installing or getting grab-site to run, please file an issue - thank you! The installation methods below are the only ones supported in our GitHub issues. Please do not modify the installation steps unless you really know what you're doing, with both Python packaging and your operating system. grab-site runs on a specific version of Python (3.7 or 3.8) and with specific dependency versions. Contents - Install on Ubuntu 18.04, 20.04, 22.04, Debian 10 (buster), Debian 11 (bullseye) ","default_branch":null,"files":null,"tree":[],"storefront":"/r/ArchiveTeam","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/ArchiveTeam/grab-site/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}