{"repo":"opendatalab/MinerU-HTML","free":true,"listed":false,"github":"https://github.com/opendatalab/MinerU-HTML","clone":"git clone https://github.com/opendatalab/MinerU-HTML.git","description":"MinerU-HTML: An SLM-powered HTML main content extractor that outputs clean HTML bodies. Perfect for Deep Research Agents, RAG applications, and training data generation.","language":"Python","stars":284,"topics":["article-extractor","corpus-tools","nlp","rag","scraping","text-extraction","trafilatura","web-scraping","webagent"],"license":"Apache-2.0","category":"scrapers-browser-automation","readme_excerpt":"MinerU-HTML English/中文 MinerU-HTML is an advanced HTML main content extraction tool based on Small Language Models (SLM). It can accurately identify and extract main content from complex web page HTML, automatically filtering out auxiliary elements such as navigation bars, advertisements, and metadata. Try on our website - Mineru-Extractor Welcome to try our online document extraction tool! Supports HTML main content extraction and OCR for various document formats. OR Download the model for local usage - MinerU-HTML-v1.1 —— Now updated to v1.1 ! News - 2026.03.19 🎉 The MinerU-HTML-v1.1 is released! A more efficient and powerful version than v1.0, now integrated with MinerU-Webkit for HTML-to-Markdown/JSON/Txt conversion. Welcome to use! - 2025.12.1 🎉 The AICC dataset is released, welcome to use! AICC dataset contains 7.3T web pages extracted and converted to Markdown format by MinerU-HTML, with cleaner main content and high-quality code, formulas, and tables. - 2025.12.1 🎉 The MinerU-HTML model is released, welcome to use! MinerU-HTML model is a fine-tuned model on Qwen3, with better performance on HTML main content extraction. - 2025.12.1 🎉 The trial page is online, welcome to visit Opendatalab-AICC to try our extraction tool! ✨ Features - 🎯 LLM-Powered Extraction : Uses state-of-the-art language models to intelligently identify main content - 📝 Extensible Output : Integrated with MinerU-Webkit to support efficient conversion of extracted HTML into Markdown, JSON and T","default_branch":null,"files":null,"tree":[],"storefront":"/r/opendatalab","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/opendatalab/MinerU-HTML/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}