{"repo":"AlibabaResearch/AdvancedLiterateMachinery","free":true,"listed":false,"github":"https://github.com/AlibabaResearch/AdvancedLiterateMachinery","clone":"git clone https://github.com/AlibabaResearch/AdvancedLiterateMachinery.git","description":"A collection of original, innovative ideas and algorithms towards Advanced Literate Machinery. This project is maintained by the OCR Team in the Language Technology Lab, Tongyi Lab, Alibaba Group.","language":"C++","stars":1834,"topics":["artificial-intelligence","documentai","multimodal","multimodal-deep-learning","ocr","computer-vision","vision-language-transformer","end-to-end-ocr","scene-text-detection","scene-text-detection-recognition"],"license":"Apache-2.0","category":"media-processing","readme_excerpt":"Advanced Literate Machinery Introduction The ultimate goal of our research is to build a system that has high-level intelligence, i.e., possessing the abilities to read, think and create , so advanced that it could even surpass human intelligence one day in the future. We name this kind of systems Advanced Literate Machinery (ALM) . To start with, we currently focus on teaching machines to read from images and documents. In years to come, we will explore the possibilities of endowing machines with the intellectual capabilities of thinking and creating , catching up with and surpassing GPT-4 and GPT-4V. This project is maintained by the 读光 OCR Team (读光-Du Guang means “ Reading The Light ”) in the Tongyi Lab, Alibaba Group. Visit our 读光-Du Guang Portal and DocMaster to experience online demos for OCR and Document Understanding. Recent Updates 2024.12 Release - CC-OCR ( CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy . paper): The CC-OCR benchmark is specifically designed for evaluating the OCR-centric capabilities of Large Multimodal Models. CC-OCR possesses a diverse range of scenarios, tasks, and challenges, which comprises four OCR-centric tracks: multi-scene text reading, multilingual text reading, document parsing, and key information extraction. It includes 39 subsets with 7,058 full annotated images, of which 41% are sourced from real applications, being released for the first time. 2024.9 Release - Platypus ( Plat","default_branch":null,"files":null,"tree":[],"storefront":"/r/AlibabaResearch","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/AlibabaResearch/AdvancedLiterateMachinery/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}