{"owner":"opendatalab","github":"https://github.com/opendatalab","claimed":false,"inventory":[],"indexed":[{"repo":"opendatalab/MinerU","github":"https://github.com/opendatalab/MinerU","description":"Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.","language":"Python","stars":77911,"topics":["extract-data","layout-analysis","ocr","parser","pdf","pdf-converter","python","document-analysis","pdf-parser","pdf-extractor-llm"],"license":null,"category":"media-processing"},{"repo":"opendatalab/MinerU-Diffusion","github":"https://github.com/opendatalab/MinerU-Diffusion","description":"[ECCV 2026] A diffusion-based framework for document OCR that replaces autoregressive decoding with block-level parallel diffusion decoding.","language":"Python","stars":614,"topics":["ai4science","diffusion","dlm","document-analysis","extract-data","layout-analysis","llada","ocr","parser","pdf"],"license":"MIT","category":"media-processing"},{"repo":"opendatalab/MinerU-HTML","github":"https://github.com/opendatalab/MinerU-HTML","description":"MinerU-HTML: An SLM-powered HTML main content extractor that outputs clean HTML bodies. Perfect for Deep Research Agents, RAG applications, and training data generation.","language":"Python","stars":284,"topics":["article-extractor","corpus-tools","nlp","rag","scraping","text-extraction","trafilatura","web-scraping","webagent"],"license":"Apache-2.0","category":"scrapers-browser-automation"}],"how_to_buy":"GET /r/opendatalab/<repo> (Accept: application/json) for any listed repo here: tree, README, price and the checkout to pay (x402; rehearse first at its test twin, simulated money). Repos under 'indexed' are free: clone them from GitHub."}