{"repo":"getomni-ai/zerox","free":true,"listed":false,"github":"https://github.com/getomni-ai/zerox","clone":"git clone https://github.com/getomni-ai/zerox.git","description":"OCR & Document Extraction using vision models","language":"TypeScript","stars":12264,"topics":["ocr","pdf"],"license":"MIT","category":"media-processing","readme_excerpt":"Zerox OCR A dead simple way of OCR-ing a document for AI ingestion. Documents are meant to be a visual representation after all. With weird layouts, tables, charts, etc. The vision models just make sense! The general logic: - Pass in a file (PDF, DOCX, image, etc.) - Convert that file into a series of images - Pass each image to GPT and ask nicely for Markdown - Aggregate the responses and return Markdown Try out the hosted version here: Or visit our full documentation at: Getting Started Zerox is available as both a Node and Python package. - Node README - npm package - Python README - pip package Feature Node.js Python ------------------------- ---------------------------- -------------------------- PDF Processing ✓ (requires graphicsmagick) ✓ (requires poppler) Image Processing ✓ ✓ OpenAI Support ✓ ✓ Azure OpenAI Support ✓ ✓ AWS Bedrock Support ✓ ✓ Google Gemini Support ✓ ✓ Vertex AI Support ✗ ✓ Data Extraction ✓ ( schema ) ✗ Per-page Extraction ✓ ( extractPerPage ) ✗ Custom System Prompts ✗ ✓ ( custom system prompt ) Maintain Format Option ✓ ( maintainFormat ) ✓ ( maintain format ) Async API ✓ ✓ Error Handling Modes ✓ ( errorMode ) ✗ Concurrent Processing ✓ ( concurrency ) ✓ ( concurrency ) Temp Directory Management ✓ ( tempDir ) ✓ ( temp dir ) Page Selection ✓ ( pagesToConvertAsImages ) ✓ ( select pages ) Orientation Correction ✓ ( correctOrientation ) ✗ Edge Trimming ✓ ( trimEdges ) ✗ Node Zerox (Node.js SDK - supports vision models from different providers like OpenAI,","default_branch":null,"files":null,"tree":[],"storefront":"/r/getomni-ai","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/getomni-ai/zerox/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}