{"repo":"Lulzx/zpdf","free":true,"listed":false,"github":"https://github.com/Lulzx/zpdf","clone":"git clone https://github.com/Lulzx/zpdf.git","description":"Zero-copy PDF text extraction library written in Zig. High-performance, memory-mapped parsing with SIMD acceleration.","language":"Zig","stars":919,"topics":["parser","pdf","simd","text-extraction","zero-copy","zig","high-performance","zero-dependency"],"license":"CC0-1.0","category":"media-processing","readme_excerpt":"zpdf (alpha stage - early version) A PDF text extraction library written in Zig. Features - Memory-mapped file reading, zero-copy where possible - Streaming text extraction with efficient arena allocation - Multiple decompression filters: FlateDecode, ASCII85, ASCIIHex, LZW, RunLength - Font encoding support: WinAnsi, MacRoman, ToUnicode CMap - XRef table and stream parsing (PDF 1.5+) - Configurable error handling (strict or permissive) - Structure tree extraction for tagged PDFs (PDF/UA) - Optional geometric reading order for non-tagged PDFs - Markdown export for structured PDFs Benchmarking Build with zig build -Doptimize=ReleaseFast , then run: This runs five zpdf extractions and, when mutool is installed, one MuPDF comparison. Treat the result as a local diagnostic rather than a controlled cross-tool benchmark: record the zpdf revision, Zig and MuPDF versions, hardware, input checksum, and run policy when publishing results. Additional corpus and accuracy tools are documented in benchmark/README.md . The full methodology uses olmOCR-Bench, veraPDF, and PDF.js corpora to keep ground-truth accuracy separate from compatibility and robustness. Initial measured findings identify concrete reading-order and dense-text gaps. Requirements - Zig 0.15.2 or later Building Usage Library CLI Python Build the shared library first: Build an installable, platform-specific wheel from the current Zig library: When developing from a checkout, the Python loader prefers ZPDF LIB and zig-out/li","default_branch":null,"files":null,"tree":[],"storefront":"/r/Lulzx","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/Lulzx/zpdf/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}