{"repo":"AKSarav/pdfstract","free":true,"listed":false,"github":"https://github.com/AKSarav/pdfstract","clone":"git clone https://github.com/AKSarav/pdfstract.git","description":"PDFStract - Extract, Chunking and Embedding Layer in Your RAG Pipeline - Available as CLI - WEBUI - API","language":"Python","stars":153,"topics":["ai","data-extraction","dataengineering","docling","ocr","pdf","pdfconversion","rag","unstructured","chunking"],"license":"Apache-2.0","category":"media-processing","readme_excerpt":"PDFStract — The Unified Data Preparation Layer for RAG Extract. Chunk. Embed in one line of code. One unified API. Switch between 10+ extraction libraries, 10+ chunking methods, and multiple embedding providers with a single parameter change. Focus on your RAG outcomes, not library dependencies. Quick Start Installation Why PDFStract? No single PDF extractor, chunker, or embedding provider works best for every document. PDFStract lets you swap, compare, and automate your data preparation strategy through a single API: - Extract : 10+ libraries (Marker, Docling, PyMuPDF4LLM, PaddleOCR, Unstructured, and more) - Chunk : 10+ methods (Token, Semantic, Sentence, Recursive, Code-aware, and more) - Embed : Multiple providers (OpenAI, Azure, Google, Ollama, Sentence Transformers) Switch any component with a single parameter change. No code refactoring needed. Python API Extract Chunk Embed Combined Pipelines CLI Web UI Open http://localhost:3000 for Web UI, http://localhost:8000 for API. What's Included Tier Libraries ------ ----------- Base pymupdf4llm, markitdown Standard + pytesseract, unstructured Advanced + marker, docling, paddleocr, deepseek, mineru Chunkers: token, sentence, semantic, recursive, code, sdpm, late, slumber, neural Embeddings: OpenAI, Azure OpenAI, Google, Ollama, Sentence Transformers, Model2Vec Documentation 📖 pdfstract.com — Full documentation, guides, and API reference Use Cases - RAG systems and knowledge bases - Document intelligence pipelines - LLM fine-","default_branch":null,"files":null,"tree":[],"storefront":"/r/AKSarav","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/AKSarav/pdfstract/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}