{"repo":"google/langextract","free":true,"listed":false,"github":"https://github.com/google/langextract","clone":"git clone https://github.com/google/langextract.git","description":"A Python library for extracting structured information from unstructured text using LLMs with precise source grounding and interactive visualization.","language":"Python","stars":38428,"topics":["llm","nlp","python","gemini-ai","information-extration","large-language-models","structured-data","gemini","gemini-api","gemini-flash"],"license":"Apache-2.0","category":"ai-agents","readme_excerpt":"LangExtract Table of Contents - Introduction - Why LangExtract? - Quick Start - Installation - API Key Setup for Cloud Models - Adding Custom Model Providers - Using OpenAI Models - Using Local LLMs with Ollama - More Examples - Romeo and Juliet Full Text Extraction - Medication Extraction - Radiology Report Structuring: RadExtract - Community Providers - Contributing - Testing - How to Cite - Disclaimer Introduction LangExtract is a Python library that uses LLMs to extract structured information from unstructured text documents based on user-defined instructions. It processes materials such as clinical notes or reports, identifying and organizing key details while ensuring the extracted data corresponds to the source text. Try the live demo &rarr; Run grounded extraction on Romeo and Juliet in your browser, no install required. Why LangExtract? 1. Precise Source Grounding: Maps every extraction to its exact location in the source text, enabling visual highlighting for easy traceability and verification. 2. Reliable Structured Outputs: Enforces a consistent output schema based on your few-shot examples, leveraging controlled generation in supported models like Gemini to guarantee robust, structured results. 3. Optimized for Long Documents: Overcomes the \"needle-in-a-haystack\" challenge of large document extraction by using an optimized strategy of text chunking, parallel processing, and multiple passes for higher recall. 4. Interactive Visualization: Instantly generates a sel","default_branch":null,"files":null,"tree":[],"storefront":"/r/google","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/google/langextract/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}