{"repo":"jztan/pdf-mcp","free":true,"listed":false,"github":"https://github.com/jztan/pdf-mcp","clone":"git clone https://github.com/jztan/pdf-mcp.git","description":"MCP server that lets Claude Code and other AI agents work through large PDFs, and whole folders of them, without overflowing context: hybrid semantic + keyword search, selective page reading, tables, images, OCR, chart data, and multi-column/CJK layouts.","language":"Python","stars":115,"topics":["ai","claude","document-processing","llm","mcp","pdf","python","codex-cli","opencode","mcp-server"],"license":"MIT","category":"mcp-servers","readme_excerpt":"pdf-mcp Surgical PDF access for AI agents: search, read, and extract without flooding context. An MCP server that lets Claude Code and other AI agents search a PDF by meaning or keyword, read only the pages that matter, and cleanly pull out tables, images, and scanned text, even from multi-column and Japanese layouts. mcp-name: io.github.jztan/pdf-mcp Try it in your browser See what your AI agent sees → Drop in any PDF, or a whole folder of them, and watch an agent triage the corpus, search across every document at once, and read only the pages that matter, using a fraction of the tokens. 100% client-side, no install required. Why pdf-mcp? Without pdf-mcp With pdf-mcp --- --- --- Large PDFs Context overflow Chunked reading Token budgeting Guess and overflow Estimated tokens before reading Finding content Load everything Hybrid search (BM25 keyword + semantic) Tables Lost in raw text Extracted and inlined per page Charts Trapped in the plot image Extracted as (x, y) data tables Multi-column PDFs Columns interleaved in extracted text Column-aware reading order ( pdf-mcp[multicolumn] ) Vertical scripts (Japanese) Columns scrambled / glyph soup Geometric reorder of vertical text (tategaki / 縦書き); CJK keyword search works on unspaced Japanese/Chinese/Korean text via a char-split FTS index Images Ignored Extracted as PNG files Repeated access Re-parse every time SQLite cache Scanned PDFs No text extracted OCR via Tesseract, parallelized across pages ( pdf read pages(ocr=True) ) Vis","default_branch":null,"files":null,"tree":[],"storefront":"/r/jztan","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/jztan/pdf-mcp/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}