{"repo":"matthsena/AlcheMark","free":true,"listed":false,"github":"https://github.com/matthsena/AlcheMark","clone":"git clone https://github.com/matthsena/AlcheMark.git","description":"Your files ready for Gen AI ✨🚀 AlcheMark is a lightweight PDF to Markdown, alchemical-inspired toolkit that transmutes PDF documents into structured Markdown pages—complete with rich metadata and named‐entity annotations—empowering you to uncover insights page by page.","language":"Python","stars":77,"topics":["markdown","pdf-converter","pdf2md"],"license":"MIT","category":"productivity","readme_excerpt":"AlcheMark. Your files ready for Gen AI ✨🚀 AlcheMark is a lightweight PDF to Markdown, alchemical-inspired toolkit that transmutes PDF documents into structured Markdown pages—complete with rich metadata and markdown element annotations—empowering you to uncover insights page by page. Installation Usage Google Colab Example Try AlcheMark AI directly in your browser with our interactive Google Colab notebook! Overview AlcheMark AI provides a seamless solution for converting PDF documents into well-structured Markdown format. The tool not only extracts the text content but also analyzes and catalogs various elements like tables, images, headings, lists, and links while tracking token counts for LLM compatibility. Key Features - PDF to Markdown Conversion : Transform PDF documents into clean, organized Markdown - Rich Metadata Extraction : Preserve document metadata including title, author, creation date - Element Analysis : Automatic detection and counting of markdown elements (headings, lists, links) - Table & Image Support : Extract and format tables and images from PDFs - Inline Image Handling : Option to keep images inline as base64 or replace with image references - Token Counting : Built-in token counting using tiktoken for LLM integration - Structured Output : Get page-by-page results with detailed metadata Extracted Data Fields Field Type Description ------- ------ ------------- metadata.file path str Path to the original PDF file metadata.page int Current page number m","default_branch":null,"files":null,"tree":[],"storefront":"/r/matthsena","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/matthsena/AlcheMark/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}