{"repo":"zyocum/reader","free":true,"listed":false,"github":"https://github.com/zyocum/reader","clone":"git clone https://github.com/zyocum/reader.git","description":"Extract clean(er), readable text from web pages via Mercury Web Parser.","language":"Python","stars":122,"topics":["mercury-parser","cleaner","web-scraping","reader","extract","readability"],"license":"MIT","category":"scrapers-browser-automation","readme_excerpt":"reader Extract clean(er), readable text from web pages via trafilatura. A note on the parser Earlier versions of this project used the Postlight Parser, which required Node.js and shelling out to its command-line driver, plus html2text for the Markdown/plain-text conversions. Both have been replaced by trafilatura, a well-maintained Python library that consistently tops content-extraction benchmarks and emits HTML, Markdown, and plain-text natively. Everything now runs in a single Python process with a single dependency. Install Clone this repository and install the dependencies with uv: Or with a classic virtual environment: Usage When wrapping markdown, lines whose markup would break if split across lines (headings, table rows, horizontal rules, and fenced code blocks) are left intact, and long tokens such as URLs are never split. Layout tables (common on older, table-based sites) are unwrapped into ordinary paragraphs so their contents read naturally, while genuine data tables are preserved; decorative tables with no text (image/spacer scaffolding) are dropped. In plain-text output, data tables are rendered as aligned text via tabulate, in any format tabulate supports ( -t/--table-format , default simple ): Rendered tables are never disturbed by -w line-wrapping, regardless of the chosen table format. The source can be a URL (fetched by trafilatura), a local HTML file, or - to read HTML from stdin — so you can also feed it pages saved locally or fetched by other tools ( cu","default_branch":null,"files":null,"tree":[],"storefront":"/r/zyocum","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/zyocum/reader/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}