{"repo":"Edgaras0x4E/paddleocr-pdf-api","free":true,"listed":false,"github":"https://github.com/Edgaras0x4E/paddleocr-pdf-api","clone":"git clone https://github.com/Edgaras0x4E/paddleocr-pdf-api.git","description":"A self-hosted PDF OCR API that converts scanned documents to markdown. Powered by PaddleOCR-VL, runs on GPU via Docker.","language":"Python","stars":24,"topics":["api","ocr","paddleocr","docker","document-ai","document-ocr","document-parsing","ocr-api","paddleocr-vl","pdf"],"license":null,"category":"media-processing","readme_excerpt":"Self-Hosted PDF OCR API for Large Documents A self-hosted PDF OCR API powered by PaddleOCR models (PaddleOCR-VL by default). Runs via Docker, processes PDFs page-by-page, and returns markdown content in JSON responses. Good support (not perfect) for Latvian and Lithuanian languages. Contents - OCR engines - Requirements - Quick start - Usage - API reference - Configuration - Database backend - Job mode - Image descriptions - API key authentication - Data persistence - Changelog OCR engines The API runs one of three OCR engines, selected with the OCR ENGINE variable. All engines share the same API and job flow: OCR ENGINE Models Output Hardware --- --- --- --- vl (default) PaddleOCR-VL-1.6 (0.9B) + PP-DocLayoutV3 Markdown, full document understanding GPU, 8.5GB VRAM text PP-OCRv5 server detection + recognition Plain text lines, no layout markup GPU or CPU structure PP-StructureV3 layout parsing (tables, formulas, reading order), no VLM Markdown GPU, 10.5GB VRAM All engines accept the same input formats: PDF, PNG, JPG, JPEG, BMP, TIFF, WEBP. Each engine has its own baked image (models included, no first-run download): Image tag Engine Runs on --- --- --- latest-vl-baked vl GPU latest-structure-baked structure GPU latest-text-baked text CPU Baked images preselect their engine, so OCR ENGINE does not need to be set. The non-baked latest image is GPU-based and works with any engine; it downloads the selected engine's models at runtime. Requirements Every image needs Docker. The GP","default_branch":null,"files":null,"tree":[],"storefront":"/r/Edgaras0x4E","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/Edgaras0x4E/paddleocr-pdf-api/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}