{"repo":"lapwat/papeer","free":true,"listed":false,"github":"https://github.com/lapwat/papeer","clone":"git clone https://github.com/lapwat/papeer.git","description":"Scrape the web in the eink era. Convert websites into ebooks and markdown.","language":"Go","stars":391,"topics":["eink","ereader","epub","mobi","kindle","remarkable","markdown","scraper","ebook","command-line"],"license":"GPL-3.0","category":"scrapers-browser-automation","readme_excerpt":"Papeer Web scraper for ereaders Features • Installation • How To Use Features Scrape websites and RSS feeds Keep relevant content only - Formatted text (bold, italic, links) - Images Save websites as Markdown, HTML, EPUB or MOBI files Use it as a an HTTP proxy Cross platform - Windows, MacOS and Linux ready Installation From source From binary Download latest release for Windows, MacOS (darwin) and Linux. MOBI support Kindle e-readers now support EPUB format Install kindlegen to export websites to Kindle compatible ebooks, Linux only. Now you can use --format=mobi in your get command. How To Use Scrape a single page The get command let's you retrieve the content of a web page. It removes ads and menus with go-readability , keeping only formatted text and images. You can chain URLs. Options Scrape a whole website recursively Display the table of contents Before scraping a whole website, it is a good idea to use the list command. This command is like a dry run , which lets you visualize the content before retrieving it . You can use several options to customize the table of contents extraction, such as selector , limit , offset , reverse and include . Type papeer list --help for more information about those options. The selector option should point to HTML tags . If you don't specify it, the selector will be automatically determined based on the links present on the page. Scrape the content Once you are satisfied with the table of contents listed by the list command, you can sc","default_branch":null,"files":null,"tree":[],"storefront":"/r/lapwat","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/lapwat/papeer/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}