{"repo":"CatchTheTornado/text-extract-api","free":true,"listed":false,"github":"https://github.com/CatchTheTornado/text-extract-api","clone":"git clone https://github.com/CatchTheTornado/text-extract-api.git","description":"Document (PDF, Word, PPTX ...) extraction and parse API using state of the art modern OCRs + Ollama supported models. Anonymize documents. Remove PII. Convert any document or picture to structured JSON or Markdown","language":"Python","stars":3170,"topics":["api","extract","json","llm","pdf","anonymization","ocr","ocr-python","pii"],"license":"MIT","category":"media-processing","readme_excerpt":"text-extract-api Convert any image, PDF or Office document to Markdown text or JSON structured document with super-high accuracy, including tabular data, numbers or math formulas. The API is built with FastAPI and uses Celery for asynchronous task processing. Redis is used for caching OCR results. Features: - No Cloud/external dependencies all you need: PyTorch based OCR (EasyOCR) + Ollama are shipped and configured via docker-compose no data is sent outside your dev/server environment, - PDF/Office to Markdown conversion with very high accuracy using different OCR strategies including llama3.2-vision, easyOCR, minicpm-v, remote URL strategies including marker-pdf - PDF/Office to JSON conversion using Ollama supported models (eg. LLama 3.1) - LLM Improving OCR results LLama is pretty good with fixing spelling and text issues in the OCR text - Removing PII This tool can be used for removing Personally Identifiable Information out of document - see examples - Distributed queue processing using Celery - Caching using Redis - the OCR results can be easily cached prior to LLM processing, - Storage Strategies switchable storage strategies (Google Drive, Local File System ...) - CLI tool for sending tasks and processing results Screenshots Converting MRI report to Markdown + JSON. Before running the example see getting started Converting Invoice to JSON and remove PII Before running the example see getting started Getting started You might want to run the app directly on your machin","default_branch":null,"files":null,"tree":[],"storefront":"/r/CatchTheTornado","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/CatchTheTornado/text-extract-api/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}