Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
A collection of original, innovative ideas and algorithms towards Advanced Literate Machinery. This project is maintained by the OCR Team in the Language Technology Lab, Tongyi Lab, Alibaba Group.
| Date | Stars |
|---|---|
| 2026-07-24 | 1833 |
| 2026-07-25 | 1833 |
| 2026-07-28 | 1832 |
| 2026-07-30 | 1832 |
| 2026-08-10 | 1834 |
| 2026-08-20 | 1835 |
| 2026-08-25 | 1834 |
| 2026-08-30 | 1834 |
| 2026-08-31 | 1835 |
| 2026-09-06 | 1835 |
| 2026-09-07 | 1836 |
| 2026-09-08 | 1835 |
| 2026-09-14 | 1834 |
| 2026-09-16 | 1835 |
| 2026-09-20 | 1835 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# Advanced Literate Machinery ## Introduction The ultimate goal of our research is to build a system that has high-level intelligence, i.e., possessing the abilities to ***read, think and create***, so advanced that it could even surpass human intelligence one day in the future. We name this kind of systems **Advanced Literate Machinery (ALM)**. To start with, we currently focus on teaching machines to ***read*** from images and documents. In years to come, we will explore the possibilities of endowing machines with the intellectual capabilities of ***thinking and creating***, catching up with and surpassing [GPT-4](https://openai.com/research/gpt-4) and [GPT-4V](https://openai.com/research/gpt-4v-system-card). This project is maintained by the **读光 OCR Team** (读光-Du Guang means “*Reading The Light*”) in the [Tongyi Lab, Alibaba Group](https://tongyi.aliyun.com/).  Visit our [读光-Du Guang Portal](https://duguang.aliyun.com/) and [DocMaster](https://www.modelscope.cn/studios/damo/DocMaster/summary) to experience online demos for OCR and Document Understanding. ## Recent Updates **2024.12 Release** - [**CC-OCR**](./Benchmarks/CC-OCR/) (*CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy*. [paper](https://arxiv.org/abs/2412.02210)): The CC-OCR benchmark is specifically designed for evaluating the OCR-centric capabilities of Large Multimodal Models. CC-OCR possesses a diverse range of scenarios, tasks, and challenges, which comprises four OCR-centric tracks: multi-scene text reading, multilingual text reading, document parsing, and key information extraction. It includes 39 subsets with 7,058 full annotated images, of which 41% are sourced from real applications, being released for the first time. **2024.9 Release** - [**Platypus**](./OCR/Platypus/) (*Platypus: A Generalized Specialist Model for Reading Text in Various Forms,* ECCV 2024. [paper](https://arxiv.org/abs/2408.14805)): Platypus introduces a novel approach to text reading from images, addressing limitations of both specialist and generalist models. Platypus leverages **a single unified architecture** to effectively recognize text in **various forms**, maintaining high accuracy and efficiency. We also introduce a **new dataset [Worms](https://www.modelscope.cn/datasets/yuekun/Worms)** which combines and partially re-labels previous datasets to support the model's development and evaluation. - [**SceneVTG**](./AIGC/SceneVTG/) (*Visual Text Generation in the Wild,* ECCV 2024. [paper](https://arxiv.org/abs/2407.14138)): We propose a visual text generator (termed SceneVTG), which can produce **high-quality text images in the wild**. Following a **two-stage paradigm**, SceneVTG leverages a Multimodal Large Language Model to recommend reasonable text regions and contents across multiple scales and levels, which are used by a conditional diffusion model as conditions to generate text images. To train SceneVTG, we also contribute a **new dataset [SceneVTG-Erase](https://www.modelscope.cn/datasets/Kpillow/SceneVTG-Erase)** with detailed OCR annotations. - [**WebRPG**](./DocumentUnderstanding/WebRPG) (*WebRPG: Automatic Web Rendering Parameters Generation for Visual Presentation,* ECCV 2024. [paper](https://arxiv.org/abs/2407.15502)): We introduce WebRPG, a novel task that focuses on **automating the generation of visual presentations** for web pages based on HTML code. In the absence of a benchmark, we created a new dataset via an **automated pipeline**. Our proposed models, built on **VAE architecture** and **custom HTML embeddings**, efficiently manage numerous web elements and rendering parameters. Comprehensive experiments, including customized quantitative evaluations, demonstrate the effectiveness of WebRPG model in generating web presentations. - [**ProcTag**](./DocumentUnderstanding/ProcTag/) (*ProcTag: Process Tagging for Assessing the Efficacy of Document Instruct
Excerpt of 11,147 characters
Read on GitHub4
Alibaba OSS · Alibaba Group Holding Limited · China
1
1
1
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:830ae0756ee6739d, topic:multimodal, topic:vision-language-model, topic:document-understanding
matched fp:830ae0756ee6739d, topic:ocr, readme:document parsing, desc:ocr