{"repo":"xberg-io/xberg","free":true,"listed":false,"github":"https://github.com/xberg-io/xberg","clone":"git clone https://github.com/xberg-io/xberg.git","description":"A polyglot document intelligence framework with a Rust core. Extract text, metadata, images, and structured data from 101 formats (115 file extensions) plus code intelligence for 371 code languages. 15 language bindings — Rust, Python, Ruby, Java, Go, PHP, Elixir, C#, TypeScript — plus CLI, REST API, and MCP server.","language":"Rust","stars":9124,"topics":["text-extraction","document-intelligence","metadata-extraction","pdf-extraction","pdfium","python","rag","table-extraction","tesseract","ffi"],"license":"MIT","category":"ai-agents","readme_excerpt":"Xberg The fast, precise document-intelligence engine — for every language. Point Xberg at anything — a PDF, a scanned image, a spreadsheet, an audio file, a URL, a whole archive, or a source tree — and get back clean text, tables, metadata, and structured data. One engine handles format detection, reading, OCR, and extraction, so you never stitch a pipeline together from a dozen libraries. 100 formats · 120 file extensions · 371 code languages · 15 language bindings · 6 output formats · OCR · transcription · embeddings The fastest, most precise open-source document and PDF-to-Markdown engine — see the benchmarks. Install · What you get · Capabilities · CLI · Docs Xberg is the next iteration of Kreuzberg. Same document-intelligence engine, rebuilt and rebranded under a fresh v1 line. --- What you get Point Xberg at anything — a PDF, a spreadsheet, a scanned image, an audio file, a URL, an archive, a source tree — and get back clean, structured content you can use right away. One core does the format detection, reading, and extraction, so you don't assemble a pipeline yourself. Call it from Rust, Python, Node.js, Go, Java, C#, Ruby, PHP, Elixir, Dart, Swift, Zig, WASM, Kotlin, or C FFI, and run it as a library, CLI tool, REST API, or MCP server. Capability What you get --- --- 100 document formats PDFs, Office, images, HTML, email, e-books, scientific publications, structured data across 120 file extensions — intelligent MIME detection, streaming for multi-GB files. URLs & the ","default_branch":null,"files":null,"tree":[],"storefront":"/r/xberg-io","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/xberg-io/xberg/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}