Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Official code implementation of Vary-toy (Small Language Model Meets with Reinforced Vision Vocabulary)
| Date | Stars |
|---|---|
| 2026-07-31 | 630 |
| 2026-08-04 | 630 |
| 2026-08-06 | 630 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
<h3><a href="">Small Language Model Meets with Reinforced Vision Vocabulary</a></h3> <a href="https://varytoy.github.io/"><img src="https://img.shields.io/badge/Project-Page-Green"></a> <a href="https://arxiv.org/abs/2401.12503"><img src="https://img.shields.io/badge/Paper-PDF-orange"></a> <a href="https://vary.xiaomy.net/"><img src="https://img.shields.io/badge/demo-blue"></a> <a href="https://zhuanlan.zhihu.com/p/679447793"><img src="https://img.shields.io/badge/zhihu-yellow"></a> <a href="https://trendshift.io/repositories/7311" target="_blank"><img src="https://trendshift.io/api/badge/repositories/7311" alt="Ucas-HaoranWei%2FVary-toy | Trendshift" style="width: 250px; height: 55px;" width="250" height="55"/></a> [Haoran Wei*](https://scholar.google.com/citations?user=J4naK0MAAAAJ&hl=en), Lingyu Kong*, Jinyue Chen, Liang Zhao, [Zheng Ge](https://joker316701882.github.io/), [En Yu](https://scholar.google.com.hk/citations?user=rWCQMNgAAAAJ&hl=zh-CN&oi=sra), [Jianjian Sun](https://scholar.google.com/citations?user=MVZrGkYAAAAJ&hl=en), Chunrui Han, [Xiangyu Zhang](https://scholar.google.com/citations?user=yuB-cfoAAAAJ&hl=en) <p align="center"> <img src="assets/vary-toy-logo.jpg" style="width: 200px" align=center> </p> <p align="center"> <a href="">Two-Stream Hypothesis for LVLMs</a> </p> ## Release - [2024/9/03] 🔥🔥🔥 We release a very strong and comprehensive OCR model [GOT-OCR2.0](https://github.com/Ucas-HaoranWei/GOT-OCR2.0). - [2024/7/21] 🎉🎉🎉 OneChart is accepted by ACM'MM 2024 **Oral**! (3.97%) - [2024/7/2] 🔥🔥🔥 Vary is accepted by ECCV2024. To thank everyone for their attention, I will release a model that performs on par with the Vary-document soon. - [2024/5/27] 🔥🔥🔥 We present a document understanding benchmark in [Fox](https://github.com/ucaslcl/Fox) . - [2024/5/24] 🔥🔥🔥 We propose a multi-page document understanding work -- [Fox](https://arxiv.org/abs/2405.14295), which supports 8-page pdf-image input !!! - [2024/4/21] 🔥🔥🔥 For OneChart, we have released the web demo in [Project Page](https://onechartt.github.io/). Have fun!! - [2024/4/21] 🔥🔥🔥 We present a Vary-tiny LAVIS codebase (for training from scratch) and the Vary-600k dataset (300K English and 300K Chinese pages) [here](https://github.com/Ucas-HaoranWei/Vary-tiny-600k) !!! - [2024/4/15]🔥🔥🔥We release a chart parsing model OneChart [here](https://github.com/LingyvKong/OneChart). - [2024/4/12]🔥🔥🔥We will release a chart parsing model based on Vary-tiny next week. The model supports both English and Chinese charts. - [2024/3/16]🔥🔥🔥I found many friends very interested in Vary-tiny(OPT-125M), so I opened source it [here](https://huggingface.co/HaoranWei/Vary-tiny-opt125M/tree/main), a PDF-dense OCR and object detection version. - [2024/1/23] 🔥Eval codes will be available soon. - [2024/1/23] 🔥🔥🔥You only need a single 1080Ti to experience all features of current LVLMs. [](https://github.com/tatsu-lab/stanford_alpaca/blob/main/LICENSE) [](https://github.com/tatsu-lab/stanford_alpaca/blob/main/DATA_LICENSE) **Usage and License Notices**: The data, code, and checkpoint are intended and licensed for research use only. They are also restricted to use that follow the license agreement of LLaMA, Vicuna, GPT-4, Qwen, and LLaVA. ## Contents - [Install](#install) - [Vary-toy Weights](#vary-weights) - [Demo](#Demo) - [Train](#train) ## Note If you have built the original [Vary](https://github.com/Ucas-HaoranWei/Vary), please rebuild this repo !!! ## Install 1. Clone this repository and navigate to the Vary folder ```bash git clone https://github.com/Ucas-HaoranWei/Vary-toy.git cd /path/to/vary-toy ``` 2. Install Package ```Shell conda create -n vary python=3.10 -y conda activate vary pip install e . ``` 3. Install Flash-Attention ``` pip install ninja pip inst
Excerpt of 6,813 characters
Read on GitHub45
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:87924ff2e101df3a, llm:Repository description: 'Official code implementation of Vary-toy (Small Language Model Meets with Reinforced Vision Vocabulary)'. Language: Python. Indicates a small language model integrated with vision vocabulary — multimodal / foundation-models / multimodal. No README provided.
matched fp:87924ff2e101df3a, llm:Repository description: 'Official code implementation of Vary-toy (Small Language Model Meets with Reinforced Vision Vocabulary)'. Language: Python. Indicates a small language model integrated with vision vocabulary — multimodal / foundation-models / multimodal. No README provided.
matched fp:87924ff2e101df3a, llm:Repository description: 'Official code implementation of Vary-toy (Small Language Model Meets with Reinforced Vision Vocabulary)'. Language: Python. Indicates a small language model integrated with vision vocabulary — multimodal / foundation-models / multimodal. No README provided.
matched fp:87924ff2e101df3a, llm:Repository description: 'Official code implementation of Vary-toy (Small Language Model Meets with Reinforced Vision Vocabulary)'. Language: Python. Indicates a small language model integrated with vision vocabulary — multimodal / foundation-models / multimodal. No README provided.