{"repo":"enoch3712/ExtractThinker","free":true,"listed":false,"github":"https://github.com/enoch3712/ExtractThinker","clone":"git clone https://github.com/enoch3712/ExtractThinker.git","description":"ExtractThinker is a Document Intelligence library for LLMs, offering ORM-style interaction for flexible and powerful document workflows.","language":"Python","stars":1592,"topics":["ai","llm","nlp","ocr","openai","python","document-image-analysis","document-intelligence","document-parsing","document-processing"],"license":"Apache-2.0","category":"ai-agents","readme_excerpt":"ExtractThinker ExtractThinker is a flexible document intelligence tool that leverages Large Language Models (LLMs) to extract and classify structured data from documents, functioning like an ORM for seamless document processing workflows. TL;DR Document Intelligence for LLMs 🚀 Key Features - Flexible Document Loaders : Support for multiple document loaders, including Tesseract OCR, Azure Form Recognizer, AWS Textract, Google Document AI, and more. - Customizable Contracts : Define custom extraction contracts using Pydantic models for precise data extraction. - Advanced Classification : Classify documents or document sections using custom classifications and strategies. - Asynchronous Processing : Utilize asynchronous processing for efficient handling of large documents. - Multi-format Support : Seamlessly work with various document formats like PDFs, images, spreadsheets, and more. - ORM-style Interaction : Interact with documents and LLMs in an ORM-like fashion for intuitive development. - Splitting Strategies : Implement lazy or eager splitting strategies to process documents page by page or as a whole. - Integration with LLMs : Easily integrate with different LLM providers like OpenAI, Anthropic, Cohere, and more. - Community-driven Development : Inspired by the LangChain ecosystem with a focus on intelligent document processing. 📦 Installation Install ExtractThinker using pip: 🛠️ Usage Basic Extraction Example Here's a quick example to get you started with ExtractThink","default_branch":null,"files":null,"tree":[],"storefront":"/r/enoch3712","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/enoch3712/ExtractThinker/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}