{"repo":"AstraBert/ingest-anything","free":true,"listed":false,"github":"https://github.com/AstraBert/ingest-anything","clone":"git clone https://github.com/AstraBert/ingest-anything.git","description":"From data to vector database effortlessly","language":"Python","stars":93,"topics":["automation","chonkie","ingestion-pipeline","llamaindex","pdf","qdrant","vector-database"],"license":"MIT","category":"databases-storage","readme_excerpt":"ingest-anything From data to vector database effortlessly ingest-anything is a python package aimed at providing a smooth solution to ingest non-PDF files into vector databases, given that most ingestion pipelines are focused on PDF/markdown files. Leveraging chonkie, PdfItDown, and LlamaIndex integrations for vector databases and data loaders, ingest-anything gives you a fully-automated pipeline for document ingestion within few lines of code! Find out more about ingest-anything on the Documentation website! (still under construction) Workflow For text files - The input files are converted into PDF by PdfItDown - The PDF text is extracted using LlamaIndex-compatible reader - The text is chunked exploiting Chonkie's functionalities - The chunks are embedded thanks to an Embedding model from Sentence Transformers, OpenAI, Cohere, Jina AI or Model2Vec - The embeddings are loaded into a LlamaIndex-compatible vector database For code files - The text is extracted from code files using LlamaIndex SimpleDirectoryReader - The text is chunked exploiting Chonkie's CodeChunker - The chunks are embedded thanks to an Embedding model from Sentence Transformers, OpenAI, Cohere, Jina AI or Model2Vec - The embeddings are loaded into a LlamaIndex-compatible vector database \\ For web data - HTML content is scraped from URLs with crawlee - HTML files are turned into PDFs with PdfItDown - The text is extracted from PDF files using LlamaIndex PyMuPdfReader - The text is chunked exploiting Chonkie","default_branch":null,"files":null,"tree":[],"storefront":"/r/AstraBert","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/AstraBert/ingest-anything/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}