{"repo":"aws-samples/layout-aware-document-processing-and-retrieval-augmented-generation","free":true,"listed":false,"github":"https://github.com/aws-samples/layout-aware-document-processing-and-retrieval-augmented-generation","clone":"git clone https://github.com/aws-samples/layout-aware-document-processing-and-retrieval-augmented-generation.git","description":"Advanced document extraction and chunking techniques for retrieval augmented generation that is aware of the layout of documents. Increases knowledge retrieval accuracy and provides control for retrieved knowledge context management","language":"Jupyter Notebook","stars":119,"topics":["information-extraction","information-retrieval","llm","retrieval-augmented-generation","vector-database","vector-search"],"license":"MIT-0","category":"ai-agents","readme_excerpt":"Layout-Aware-Document-Extraction-Chunking-and-Indexing Prerequisite: - Amazon Bedrock model access - Deploy Embedding and Text Generation Large Language Models with SageMaker JumpStart - Create OpenSearch Domian. This solution uses IAM as master user for fine-grained access control. This example notebook guides you through the process of utilizing Amazon Textract's layout feature. This feature allows you to extract content from your document while maintaining its layout and reading format. Amazon Textract Layout feature is able to detect the following sections: - Titles - Headers - Sub-headers - Text - Tables - Figures - List - Footers - Page Numbers - Key-Value pairs Here is a snippet of Textract Layout feature on a page of Amazon Sustainability report using the Textract Console UI: The Amazon Textract Textractor Library is a library that seamlessly works with Textract features to aid in document processing. You can start by checking out the examples in the documentation. This notebook utilizes the Textractor library to interact with Amazon Textract and interpret its response. It enriches the extracted document text with XML tags to delineate sections, facilitating layout-aware chunking and document indexing into a Vector Database (DB). This process aims to enhance Retrieval Augmented Generation (RAG) performance. DOCUMENT PROCESSING AND INDEXING 1. Upload multi-page document to Amazon S3. 2. Call Amazon Textract Start Document Analysis api call to extract Document Text incl","default_branch":null,"files":null,"tree":[],"storefront":"/r/aws-samples","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/aws-samples/layout-aware-document-processing-and-retrieval-augmented-generation/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}