{"repo":"NameetP/pdfmux","free":true,"listed":false,"github":"https://github.com/NameetP/pdfmux","clone":"git clone https://github.com/NameetP/pdfmux.git","description":"PDF extraction that audits its own output — and certifies any other extractor's, catching pages they silently dropped. Verify signed manifests offline: free, MIT, no account. 0.903 on opendataloader-bench, #2 of 8 engines. 7-tool MCP server.","language":"Python","stars":81,"topics":["llm","mcp","ocr","pdf","pdf-to-markdown","python","pdf-to-json","structured-extraction","ai-agent","docling"],"license":"MIT","category":"media-processing","readme_excerpt":"pdfmux Self-healing PDF extraction that flags the pages it can't read instead of dropping them — and now certifies any extractor's output for silent drops. Open-source LlamaParse alternative for RAG pipelines, MCP server for Claude Desktop, LangChain + LlamaIndex loaders. pdfmux extracts PDFs and checks its own work — and now certifies any extractor's, telling you which pages it silently dropped. Free, MIT. Patent-pending method. pip install pdfmux . Two jobs, one tool: - Self-healing extraction. The only PDF extractor that audits its own output. Catches blank pages, scrambled columns, broken tables — re-extracts them with a stronger backend, and flags what it still can't read instead of silently dropping it. So your LLM gets clean data, not silent garbage. Routes each page to the best of 7 built-in extraction backends + BYOK LLM fallback (Gemini / Claude / GPT-4o / Ollama). One CLI. One API. Zero config. - Certify Anything — new in v1.8.1. pdfmux verify audits any extraction engine's output against the source PDF — Reducto, Mistral OCR, LlamaParse, Docling, your in-house parser — and tells you which pages it silently dropped. Free, MIT, patent-clean. Install That handles digital PDFs. For any real-world batch, install pdfmux[ocr] too — almost every directory of PDFs has at least one scan, and without OCR those pages return empty text: Other backends, by document type: Requires Python 3.11+. Quick Start CLI Python For batch processing, use batch extract() — not a subprocess.r","default_branch":null,"files":null,"tree":[],"storefront":"/r/NameetP","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/NameetP/pdfmux/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}