{"repo":"zai-org/GLM-OCR","free":true,"listed":false,"github":"https://github.com/zai-org/GLM-OCR","clone":"git clone https://github.com/zai-org/GLM-OCR.git","description":"GLM-OCR: Accurate × Fast × Comprehensive","language":"Python","stars":7290,"topics":["glm","image2text","ocr"],"license":"Apache-2.0","category":"media-processing","readme_excerpt":"GLM-OCR 中文阅读 👋 Join our WeChat and Discord community 📖 Check out the GLM-OCR technical report 📍 Use GLM-OCR's API Model Introduction GLM-OCR is a multimodal OCR model for complex document understanding, built on the GLM-V encoder–decoder architecture. It introduces Multi-Token Prediction (MTP) loss and stable full-task reinforcement learning to improve training efficiency, recognition accuracy, and generalization. The model integrates the CogViT visual encoder pre-trained on large-scale image–text data, a lightweight cross-modal connector with efficient token downsampling, and a GLM-0.5B language decoder. Combined with a two-stage pipeline of layout analysis and parallel recognition based on PP-DocLayout-V3, GLM-OCR delivers robust and high-quality OCR performance across diverse document layouts. Key Features - State-of-the-Art Performance : Achieves a score of 94.62 on OmniDocBench V1.5, ranking #1 overall, and delivers state-of-the-art results across major document understanding benchmarks, including formula recognition, table recognition, and information extraction. - Optimized for Real-World Scenarios : Designed and optimized for practical business use cases, maintaining robust performance on complex tables, code-heavy documents, seals, and other challenging real-world layouts. - Efficient Inference : With only 0.9B parameters, GLM-OCR supports deployment via vLLM, SGLang, and Ollama, significantly reducing inference latency and compute cost, making it ideal for high-c","default_branch":null,"files":null,"tree":[],"storefront":"/r/zai-org","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/zai-org/GLM-OCR/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}