{"repo":"EvilFreelancer/img2md-vlm-ocr","free":true,"listed":false,"github":"https://github.com/EvilFreelancer/img2md-vlm-ocr","clone":"git clone https://github.com/EvilFreelancer/img2md-vlm-ocr.git","description":"A comprehensive service for extracting document structure and content from images using advanced computer vision and vision-language models (VLM)","language":"JavaScript","stars":11,"topics":["api","fastapi","image","markdown","md","ocr","react","typescript","vlm"],"license":"MIT","category":"media-processing","readme_excerpt":"img2md VLM OCR Overview img2md VLM OCR is a comprehensive service for extracting document structure and content from images using advanced computer vision and vision-language models (VLM). The system combines YOLO-based document layout segmentation with OpenAI-compatible VLM models to accurately detect, classify, and extract text content from document images, converting them to structured Markdown format. Key Features - Document Layout Segmentation : Uses YOLOv8-based models to detect and classify document elements (text blocks, tables, images, headers, etc.) - Vision-Language Model Integration : Leverages OpenAI-compatible VLM models (default: Qwen2.5-VL) for intelligent text extraction - REST API : FastAPI-based backend with comprehensive error handling and logging - Web Interface : Modern React-based UI with drag-and-drop upload, real-time preview, and result visualization - CLI Tools : Command-line utilities for batch PDF processing and document analysis - Docker Support : Complete containerization with Docker Compose for easy deployment Architecture Backend Components - FastAPI Server : RESTful API with CORS support and automatic OpenAPI documentation - Segmentation Service : YOLO-based document layout detection using pre-trained models from Hugging Face - VLM Service : Vision-language model integration for intelligent text extraction and Markdown conversion - OpenAI Service : Configurable API client supporting custom endpoints and proxy configurations Frontend Component","default_branch":null,"files":null,"tree":[],"storefront":"/r/EvilFreelancer","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/EvilFreelancer/img2md-vlm-ocr/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}