{"repo":"genieincodebottle/parsemypdf","free":true,"listed":false,"github":"https://github.com/genieincodebottle/parsemypdf","clone":"git clone https://github.com/genieincodebottle/parsemypdf.git","description":"Collection of PDF parsing libraries like AI based docling, claude, openai, gemini, meta's llama-vision, unstructured-io, and pdfminer, pymupdf, pdfplumber etc for efficient snapshot, text, table, and metadata extraction.","language":"Python","stars":197,"topics":["camelot","claude","docling","llama-parse","markitdown","openai","pymupdf","pypdf","unstructured-io","llama-vision"],"license":"MIT","category":"ai-agents","readme_excerpt":"&nbsp; &nbsp; &nbsp; &nbsp; 👉 GenAI Roadmap - 2025 🖼️ OCR with Multimodal Vision Language Models 📑 Complex PDF Parsing Comprehensive example code for extracting content from complex PDFs with mixed elements, including text and image data extraction. Includes two Streamlit apps : 1. PDF Parser & RAG Evaluator ( pdf parser app.py ) - Parse PDFs with 13 different parsers + ask questions using RAG 2. VLM OCR App ( vlm ocr app.py ) - Extract text from images using Vision Language Models (Claude, Gemini, GPT-4o, Mistral-OCR, Ollama, OmniAI) Also, check - PDF Parsing Guide 🎥 YouTube Video: Walkthrough on setup and running the app 📦 Implementation Options 1. ☁️ Paid - API Based Methods Model Provider Models Details Example Code Doc -------------- ------- --------- :------------: :---: Anthropic claude-opus-4-20250514 , claude-sonnet-4-20250514 , claude-3-7-sonnet-20250219 , claude-3-5-sonnet-20241022 Claude 4/3.7/3.5 Sonnet is a multimodal AI model developed by Anthropic, capable of processing both text and images. It excels in visual reasoning tasks, such as interpreting charts and graphs, and can accurately transcribe text from imperfect images. Supports native PDF input via base64 encoding. Code Doc Gemini gemini-2.5-pro , gemini-2.5-flash , gemini-2.5-flash-lite-preview-06-17 , gemini-2.0-flash , gemini-2.0-flash-lite Gemini 2.5/2.0 models offer superior speed, native tool integration, and multimodal generation capabilities. Support 1M token context window, native PDF input,","default_branch":null,"files":null,"tree":[],"storefront":"/r/genieincodebottle","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/genieincodebottle/parsemypdf/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}