{"repo":"luisleo526/doc2mark","free":true,"listed":false,"github":"https://github.com/luisleo526/doc2mark","clone":"git clone https://github.com/luisleo526/doc2mark.git","description":"AI-powered Python library that converts any document (PDF, Word, Excel, PowerPoint, HTML) to clean Markdown while preserving complex tables and layouts using AI-Powered OCR technology.","language":"Python","stars":52,"topics":["complex-table-structure","langchain","markdown","ocr","openai","rag","python"],"license":"MIT","category":"ai-agents","readme_excerpt":"doc2mark Turn any document into clean, RAG-ready Markdown — in one line. doc2mark converts PDFs, Office files, images, HTML, and more into Markdown that's faithful to the original — merged table cells survive, headers/footers get stripped, and scanned pages are read by a vision LLM into a structured schema instead of a flat text blob. It's built for the part everyone hits after conversion: feeding clean, structured text to an LLM or a retrieval pipeline. --- Why doc2mark Most \"doc → markdown\" tools are fine until the document gets real — a financial statement with a merged-cell header, a scanned invoice, a slide deck with a chart. doc2mark is built for exactly those: - 🧩 Complex tables survive. Merged cells ( rowspan / colspan ), multi-level headers, and group headers are preserved as clean HTML — not flattened into a mangled markdown grid. (See below.) - 🧠 Structured OCR, not a text dump. Scanned/image pages return an OCRPage with a hard wall between verbatim transcription and model interpretation (document type, summary, key-value fields, figures, entities). - 🧹 Noise removed. Repeated page headers, footers, and page numbers are detected and dropped before they pollute your markdown or your RAG chunks. - 🔌 Bring your own model. OpenAI, Google Gemini (Vertex AI), local Tesseract, or any OpenAI-compatible endpoint (Ollama, vLLM, LM Studio) via base url . - 🪶 No ML stack to host. No multi-gigabyte model downloads — text parsing is local and deterministic; OCR calls a host","default_branch":null,"files":null,"tree":[],"storefront":"/r/luisleo526","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/luisleo526/doc2mark/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}