{"repo":"PaddlePaddle/PaddleOCR","free":true,"listed":false,"github":"https://github.com/PaddlePaddle/PaddleOCR","clone":"git clone https://github.com/PaddlePaddle/PaddleOCR.git","description":"Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.","language":"Python","stars":87875,"topics":["ocr","chineseocr","pdf2markdown","pp-ocr","pp-structure","document-parsing","document-translation","kie","ai4science","pdf-extractor-rag"],"license":"Apache-2.0","category":"media-processing","readme_excerpt":"Global Leading OCR Toolkit & Document AI Engine English 简体中文 繁體中文 日本語 한국어 Français Русский Español العربية PaddleOCR converts PDF documents and images into structured, LLM-ready data (JSON/Markdown) with industry-leading accuracy. With 70k+ Stars and trusted by top-tier projects like Dify, RAGFlow, and Cherry Studio, PaddleOCR is the bedrock for building intelligent RAG and Agentic applications. 🚀 Key Features 📄 Intelligent Document Parsing (LLM-Ready) Transforming messy visuals into structured data for the LLM era. SOTA Document VLM : Featuring PaddleOCR-VL-1.6 (0.9B) , the industry's leading lightweight vision-language model for document parsing. It achieves 96.3% accuracy on OmniDocBench v1.6, leads in text, formula, and table recognition, and shows significantly enhanced capabilities in ancient documents, rare characters, seals, and charts, with structured outputs in Markdown and JSON formats. Structure-Aware Conversion : Powered by PP-StructureV3 , seamlessly convert complex PDFs and images into Markdown or JSON . Unlike the PaddleOCR-VL series models, it provides more fine-grained coordinate information, including table cell coordinates, text coordinates, and more. Production-Ready Efficiency : Achieve commercial-grade accuracy with an ultra-small footprint. Outperforms numerous closed-source solutions in public benchmarks while remaining resource-efficient for edge/cloud deployment. 🔍 Universal Text Recognition (Scene OCR) The global gold standard for high-speed, mu","default_branch":null,"files":null,"tree":[],"storefront":"/r/PaddlePaddle","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/PaddlePaddle/PaddleOCR/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}