Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model
| Date | Stars |
|---|---|
| 2026-07-31 | 327 |
| 2026-08-04 | 327 |
| 2026-08-06 | 327 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
<div align="center"> <br> <h3>VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model</h3> Xianwei Zhuang<sup>1*</sup> Yuxin Xie<sup>1*</sup> Yufan Deng<sup>1*</sup> <br> Liming Liang<sup>1</sup> Jinghan Ru<sup>1</sup> Yuguo Yin<sup>1</sup> Yuexian Zou <sup>1</sup> <sup>1</sup> Peking University [](https://arxiv.org/pdf/2501.12327) [](https://vargpt-1.github.io/) [](https://huggingface.co/VARGPT-family/VARGPT_LLaVA-v1) [](https://huggingface.co/datasets/VARGPT-family/VARGPT_datasets) [](https://github.com/VARGPT-family/VARGPT/blob/main/LICENSE) [](https://mp.weixin.qq.com/s/mXFn-9QwU9pO1HWaz8uV0A) </div> ## News * **[2025-04-07]** The technical report for VARGPT-v1.1 is released at https://arxiv.org/pdf/2504.02949. * **[2025-04-01]** We release a more powerful model **VARGPT-v1.1** and all the training code of **VARGPT** and **VARGPT-v1.1** at https://github.com/VARGPT-family/VARGPT-v1.1. For the convenience of maintenance, we have closed the issue of this repository and will mainly update and maintain it in [VARGPT-v1.1](https://github.com/VARGPT-family/VARGPT-v1.1) and [VARGPT-v1.1 issues](https://github.com/VARGPT-family/VARGPT-v1.1/issues). 🔥🔥🔥🔥🔥🔥🔥🔥 * **[2025-01-22]** We release the datasets for training VARGPT (**7B+2B**). 🔥🔥 * **[2025-01-21]** We release the **model and inference code** of VARGPT (**7B+2B**) for multimodal understanding and generation including image captioning, visual question answering (VQA), text-to-image generation. 🔥🔥 ## What is the new about VARGPT? <p align="center"> <img src="docs/understanding_vis.png" width="333"> </p> <p align="center"> <img src="docs/generation_vis.png" width="666"> </p> Below is a comparison among understanding only, generation only, and unified (understanding \& generation) models. `Image` and `Text` indicate the representations from specific input modalities. VARGPT modeling understanding and generation as two distinct paradigms within a unified model: **predicting the next token for visual understanding and predicting the next scale for visual generation** <p align="center"> <img src="docs/comparison.png" width="666"> </p> Below is an overview of **VARGPT** and proposed unified three-stage instruction tuning. VARGPT is capable of handling image captioning, visual question answering, text-to-image generation and mixed modality generation. <p align="center"> <img src="docs/vargpt_methods.png" width="666"> </p> <p align="center"> <img src="docs/vargpt_training.png" width="666"> </p> The visual generation capabilities of VARGPT are currently constrained by the limitations of its training data, which is derived from the ImageNet dataset (1.28M images). In the forthcoming iteration, we intend to implement substantial enhancements to both the quality and quantity of the training data. <br/> ## TODO - [X] Release the inference code. - [X] Release the code for evaluation. - [X] Release the model checkpoint. - [X] Release the datasets. - [X] Supporting stronger visual generation capabilities. - [X] Release the training code at [VARGPT-v1.1](https://github.com/VARGPT-family/VARGPT-v1.1) ## Hugging Face models and annotations The VARGPT checkpoints can be found on [Hugging Face](https://huggingface.co): * [VARGPT-family/VARGPT_LLaVA-v1](https://huggingface.co/VARGPT-family/VARGPT_LLaVA-v1) The instruction for training data can be found on [Hugging Face](https://huggingface.co): * [VARGPT-family/VARGPT_datasets](https://hugg
Excerpt of 13,789 characters
Read on GitHubWould you bet a product on this? Bounded 0–100 and slow moving.
matched fp:bdf6e605f6421630, desc:multimodal