{"repo":"openzim/zimit","free":true,"listed":false,"github":"https://github.com/openzim/zimit","clone":"git clone https://github.com/openzim/zimit.git","description":"Make a ZIM file from any Web site and surf offline!","language":"Python","stars":830,"topics":["zim","docker","webscraping","scraper"],"license":"GPL-3.0","category":"scrapers-browser-automation","readme_excerpt":"Zimit ===== Zimit is a scraper allowing to create ZIM file) from any Web site. Zimit adheres to openZIM's Contribution Guidelines. Zimit has implemented openZIM's Python bootstrap, conventions and policies v2.0.0 . Capabilities and known limitations -------------------- While we would like to support as many websites as possible, making an offline archive of any website with a versatile tool obviously has some limitations. Most capabilities and known limitations are documented in warc2zim README. There are also some limitations in Browsertrix Crawler (used to fetch the website) and wombat (used to properly replay dynamic web requests), but these are not (yet?) clearly documented. Technical background -------------------- Zimit runs a fully automated browser-based crawl of a website property and produces a ZIM of the crawled content. Zimit runs in a Docker container. The system: - runs a website crawl with Browsertrix Crawler, which produces WARC files - converts the crawled WARC files to a single ZIM using warc2zim The zimit.py is the entrypoint for the system. After the crawl is done, warc2zim is used to write a zim to the /output directory, which should be mounted as a volume to not loose the ZIM created when container stops. Using the --keep flag, the crawled WARCs and few other artifacts will also be kept in a temp directory inside /output Usage ----- zimit is intended to be run in Docker. Docker image is published at https://github.com/orgs/openzim/packages/container/pac","default_branch":null,"files":null,"tree":[],"storefront":"/r/openzim","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/openzim/zimit/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}