{"repo":"bzsanti/oxidizePdf","free":true,"listed":false,"github":"https://github.com/bzsanti/oxidizePdf","clone":"git clone https://github.com/bzsanti/oxidizePdf.git","description":"Pure Rust PDF library for AI/RAG: structure-aware chunking, no ML, no C deps.","language":"Rust","stars":185,"topics":["data-extraction","document-processing","pdf","pdf-generation","pdf-parser","rust","text-extraction","digital-signatures","encryption","invoice"],"license":"MIT","category":"security-tools","readme_excerpt":"oxidize-pdf The Rust PDF library built for AI. Parse any PDF into structure-aware, embedding-ready chunks with one line of code. Pure Rust, zero C dependencies, 99.3% success rate on 9,000+ real-world PDFs. Why oxidize-pdf for RAG? Most PDF libraries give you a wall of text. oxidize-pdf gives you structured, metadata-rich chunks ready for your vector store: What you get Why it matters --- --- chunk.full text Heading context prepended -- better embeddings chunk.page numbers Citation back to source pages chunk.bounding boxes Spatial position for visual grounding chunk.element types Filter by \"table\", \"title\", \"paragraph\" chunk.token estimate Right-size chunks for your model's context window chunk.heading context Section awareness without post-processing Performance : Pure Rust, 3,000-4,000 pages/sec generation, 85ms full-text extraction for a 930KB PDF. Quick Start RAG Pipeline -- One Liner Custom Chunk Size Contextual Retrieval (no-ML) Prepend a deterministic context snippet to each chunk's full text before embedding — the no-ML analogue of Contextual Retrieval / Late Chunking, which cut retrieval failures substantially by situating each chunk in its document. No LLM, no GPU: the prefix is a pure function of the document metadata and the chunk's heading breadcrumb, so it is fully reproducible (stable chunk id ). The display text field stays context-free. ContextFormat::Prose renders the same facts as one natural-language sentence. Modes: ContextMode::None (bare text), Heading ","default_branch":null,"files":null,"tree":[],"storefront":"/r/bzsanti","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/bzsanti/oxidizePdf/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}