{"repo":"NanoNets/docstrange","free":true,"listed":false,"github":"https://github.com/NanoNets/docstrange","clone":"git clone https://github.com/NanoNets/docstrange.git","description":"Extract and convert data from any document, images, pdfs, word doc, ppt or URL into multiple formats (Markdown, JSON, CSV, HTML) with intelligent structured data extraction and advanced OCR.","language":"Python","stars":1526,"topics":["image-to-markdown","llm","markdown","ocr","pdf-to-markdown","structured-data","ai","document-parser","document-parsing","pdf-parser"],"license":"MIT","category":"media-processing","readme_excerpt":"DocStrange 🚀 Try DocStrange Online → DocStrange DocStrange converts documents to Markdown, JSON, CSV, and HTML quickly and accurately. - Converts PDF, image, PPTX, DOCX, XLSX, and URL files. - Formats tables into clean, LLM-optimized Markdown. - Powered by an upgraded 7B model for higher accuracy and deeper document understanding. - Extracts text from images and scanned documents with advanced OCR. - Removes page artifacts for clean, readable output. - Does structured extraction, given specific fields or a JSON schema. - Includes a built-in, local Web UI for easy drag-and-drop conversion. - Offers a free cloud API for instant processing or a 100% private, local mode. - Works on GPU or CPU when running locally. - Integrates with Claude Desktop via an MCP server for intelligent document navigation. --- Processing Modes ☁️ Free Cloud Processing upto 10000 docs per month ! Extract documents data instantly with the cloud processing - no complex setup needed 🔒 Local Processing ! Use gpu mode for 100% local processing - no data sent anywhere, everything stays on your machine. What's New August 2025 - 🚀 Major Model Upgrade : The core model has been upgraded to 7B parameters , delivering significantly higher accuracy and deeper understanding of complex documents. - 🖥️ Local Web Interface : Introducing a built-in, local GUI. Now you can convert documents with a simple drag-and-drop interface, 100% offline. --- About Convert and extract data from PDF, DOCX, images, and more into cle","default_branch":null,"files":null,"tree":[],"storefront":"/r/NanoNets","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/NanoNets/docstrange/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}