{"repo":"zytedata/web-snap","free":true,"listed":false,"github":"https://github.com/zytedata/web-snap","clone":"git clone https://github.com/zytedata/web-snap.git","description":"Create \"perfect\" snapshots of web pages","language":"JavaScript","stars":34,"topics":["javascript","web-archives","web-archiving","playwright","capture-page"],"license":"MIT","category":"scrapers-browser-automation","readme_excerpt":"Web-snaphots Create \"perfect\" snapshots of web pages. Install Usage This will open a Chrome-like browser, show you the page and create an output file called by default: \"snapshot en.wikipedia.org.json\" To restore this snapshot file, you can use: This will open a Chrome-like browser, show the page and you can read it even if you're offline. You can also save and restore more complicated pages, like Amazon products: Note that some pages should be scrolled a little bit and hover some elements, to make sure all the page and images are loaded before the snapshot is taken. This is not a limitation of web-snap, it's how modern browsers and pages are intentionally built to load resources lazily, on demand. For a complete example, with all the flags: This will store the page just like before, but it will do a lot of pre-processing, to reduce the snapshot size from 1.3MB , to only 27K (48x smaller), without losing any useful information. The --gzip flag will archive the JSON using GZIP. It is totally safe to use. The --rm flag, or --removeElems , will remove the specified page elements, using selectors. This can be used to remove useless elements so you can focus on the important content and reduce the snapshot size. The --css flag, or --addCSS , will add custom CSS on the page, before creating the snapshot. This can be used to change the font size, or move some elements to make the page look nicer. The --drop , or --dropRequests flag, will drop all HTTP requests matching, with regex. ","default_branch":null,"files":null,"tree":[],"storefront":"/r/zytedata","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/zytedata/web-snap/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}