{"repo":"NanoNets/docext","free":true,"listed":false,"github":"https://github.com/NanoNets/docext","clone":"git clone https://github.com/NanoNets/docext.git","description":"An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)","language":"Python","stars":2066,"topics":["document","document-analysis","extraction","llms","machine-learning","nlp","ocr","rag","unstructured-data","vlms"],"license":"Apache-2.0","category":"machine-learning","readme_excerpt":"docext An on-premises document information extraction and benchmarking toolkit. New Model Release: Nanonets-OCR-s We're excited to announce the release of Nanonets-OCR-s, a compact 3B parameter model specifically trained for efficient image to markdown conversion with semantic understanding for images, signatures, watermarks, etc.! 📢 Read the full announcement 🤗 Hugging Face model Overview docext is a comprehensive on-premises document intelligence toolkit powered by vision-language models (VLMs). It provides three core capabilities: 📄 PDF & Image to Markdown Conversion : Transform documents into structured markdown with intelligent content recognition, including LaTeX equations, signatures, watermarks, tables, and semantic tagging. 🔍 Document Information Extraction : OCR-free extraction of structured information (fields, tables, etc.) from documents such as invoices, passports, and other document types, with confidence scoring. 📊 Intelligent Document Processing Leaderboard : A comprehensive benchmarking platform that tracks and evaluates vision-language model performance across OCR, Key Information Extraction (KIE), document classification, table extraction, and other intelligent document processing tasks. Features PDF and Image to Markdown Convert both PDF and images to markdown with content recognition and semantic tagging. - LaTeX Equation Recognition : Convert both inline and block LaTeX equations in images to markdown. - Intelligent Image Description : Generate a d","default_branch":null,"files":null,"tree":[],"storefront":"/r/NanoNets","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/NanoNets/docext/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}