{"repo":"xberg-io/html-to-markdown","free":true,"listed":false,"github":"https://github.com/xberg-io/html-to-markdown","clone":"git clone https://github.com/xberg-io/html-to-markdown.git","description":"High performance and CommonMark compliant HTML to Markdown converter. Maintained by the Kreuzberg team. Kreuzberg is a fast, polyglot document intelligence engine with a Rust core. It extracts structured data from 98+ document formats using streaming parsers and built-in OCR.","language":"HTML","stars":851,"topics":["html-converter","markdown-converter","rag","text-extraction","text-processing","hocr","html","markdown"],"license":"MIT","category":"ai-agents","readme_excerpt":"html-to-markdown Turn messy, real-world HTML into clean Markdown — from the language you already work in. What and Why? Feed html-to-markdown the HTML you actually have — unclosed tags, CDATA, custom elements, broken entities, nested tables, mixed encodings — and get back clean CommonMark (or Djot) without losing content. One convert() call does it, and it returns the same result whether you run it from Python, TypeScript, Go, Ruby, Java, or 11 more languages. You get more than the text: pull page metadata (Open Graph, Twitter, JSON-LD) and structured tables in the same pass, or hook into the conversion to reshape the output. It is fast enough for whole-corpus jobs, and the messy-input handling is automatic — you never choose a parsing strategy or tune anything to get correct output. Features Feature Description ------- ----------- 16 languages, one Rust core Rust, Python, Node.js, WASM, Java, Go, C#, PHP, Ruby, Elixir, R, Dart, Kotlin (Android), Swift, Zig, and a C ABI Tiered dispatch Byte scanner → DOM walker → html5ever repair, with byte-equal output across tiers Real-HTML robust Unclosed tags, CDATA, custom elements, malformed entities, nested tables, mixed encodings — handled without losing content GFM tables Padded cells, alignment, and pipe escaping Djot output Set output format = \"djot\" to emit Djot instead of Markdown Metadata extraction Parse into structured metadata (Open Graph, Twitter, JSON-LD, microdata, RDFa, header hierarchy) Inline images Opt-in mirroring of ","default_branch":null,"files":null,"tree":[],"storefront":"/r/xberg-io","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/xberg-io/html-to-markdown/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}