Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Research Trends in LLM-guided Multimodal Learning.
| Date | Stars |
|---|---|
| 2026-07-31 | 355 |
| 2026-08-03 | 355 |
| 2026-08-06 | 355 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# Awesome-Multimodal-LLM Research Trends in `LLM-guided Multimodal` Learning. - Multi Modalities: - text, vision (image and video), audio, ... - Large Language Model (LLM) Backbones: - LLaMA, Alpaca, Vicuna, Bloom, GLM, OPT, ... - LLM should be `open-source and research-friendly` - relatively small backbones (e.g., BART and T5) are also OK - Learning Techniques: - full fine-tuning, parameter-efficient tuning (Adapter, LoRA, ... ) - in-context learning, instruction tuning - ... - Examples of `LLM-guided Multimodal Model` : - OpenFlamingo, MiniGPT-4, Otter, InstructBILP, BLIVA ... - Examples of `Evaluation on Multimodal LLM`: - MultiInstruct, POPE, AttackVLM, ... ## 2023 August - **BLIVA: A Simple Multimodal LLM for Better Handling of Text-Rich Visual Questions**. arXiv:2308.09936. *Wenbo Hu, Yifan Xu, Yi Li, Weiyue Li, Zeyuan Chen, Zhuowen Tu.* [[Paper](https://arxiv.org/abs/2308.09936)] [[Code](https://github.com/mlpc-ucsd/BLIVA)] `Backbone`: Vicuna-7B and Flan-T5-XXL (11B). ## 2023 June - **LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day**. arXiv:2306.00890. *Chunyuan Li, Cliff Wong, Sheng Zhang, Naoto Usuyama, Haotian Liu, Jianwei Yang, Tristan Naumann, Hoifung Poon, Jianfeng Gao.* [[Paper](https://arxiv.org/abs/2306.00890)] [[Code](https://github.com/microsoft/LLaVA-Med)] `Backbone`: based on LLaVA (with Vicuna-13B). - **Ziya-Visual**. *Jiaxing Zhang and Ruyi Gan and Junjie Wang and Yuxiang Zhang and Lin Zhang and Ping Yang and Xinyu Gao and Ziwei Wu and Xiaoqun Dong and Junqing He and Jianheng Zhuo and Qi Yang and Yongfeng Huang and Xiayu Li and Yanghan Wu and Junyu Lu and Xinyu Zhu and Weifeng Chen and Ting Han and Kunhao Pan and Rui Wang and Hao Wang and Xiaojun Wu and Zhongshen Zeng and Chongpei Chen.* [Paper] [[Code](https://huggingface.co/IDEA-CCNL/Ziya-BLIP2-14B-Visual-v1)] `Backbone`: based on Ziya-LLaMA-13B-v1. - **Video-LLaMA: An Instruction-Finetuned Visual Language Model for Video Understanding**. arXiv:2306.02858. *Zhang, Hang and Li, Xin and Bing, Lidong.* [[Paper](https://arxiv.org/abs/2306.02858)] [[Code](https://github.com/DAMO-NLP-SG/Video-LLaMA/tree/main)] `Backbone`: Vicuna-7B and Vicuna-13B. ## 2023 May - **Transfer Visual Prompt Generator across LLMs**. arXiv:2305.01278. *Ao Zhang, Hao Fei, Yuan Yao, Wei Ji, Li Li, Zhiyuan Liu, Tat-Seng Chua.* [[Paper](https://arxiv.org/abs/2305.01278)] [[Code](https://github.com/VPGTrans/VPGTrans)] `Backbone`: OPT (125M, 350M, 1.3B, and 2.7B) and Flan-T5 (base, large, and XL). - **LMEye: An Interactive Perception Network for Large Language Models**. arXiv:2305.03701. *Yunxin Li, Baotian Hu, Xinyu Chen, Lin Ma, Min Zhang.* [[Paper](https://arxiv.org/abs/2305.03701)] [[Code](https://github.com/YunxinLi/LingCloud)] `Backbone`: LLaMA-7B, LLaMA-13B and Bloomz-7B. - **Otter: A Multi-Modal Model with In-Context Instruction Tuning**. arXiv:2305.03726. *Bo Li, Yuanhan Zhang, Liangyu Chen, Jinghao Wang, Jingkang Yang, Ziwei Liu.* [[Paper](https://arxiv.org/abs/2305.03726)] [[Code](https://github.com/Luodian/Otter)] `Backbone`: based on OpenFlamingo-9B. - **X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages**. arXiv:2305.04160. *Feilong Chen, Minglun Han, Haozhi Zhao, Qingyang Zhang, Jing Shi, Shuang Xu, Bo Xu.* [[Paper](https://arxiv.org/abs/2305.04160)] [[Code](https://github.com/phellonchen/X-LLM)] `Backbone`: ChatGLM. - **MultiModal-GPT: A Vision and Language Model for Dialogue with Humans**. arXiv:2305.04790. *Tao Gong, Chengqi Lyu, Shilong Zhang, Yudong Wang, Miao Zheng, Qian Zhao, Kuikun Liu, Wenwei Zhang, Ping Luo, Kai Chen.* [[Paper](https://arxiv.org/abs/2305.04790)] [[Code](https://github.com/open-mmlab/Multimodal-GPT)] `Backbone`: based on OpenFlamingo. - **VideoChat: Chat-Centric Video Understanding**. arXiv:2305.06355. *KunChang Li, Yinan He, Yi Wang, Yizhuo Li, Wenhai Wang, Ping Luo, Y
Excerpt of 14,680 characters
Read on GitHubWould you bet a product on this? Bounded 0–100 and slow moving.
matched fp:a8c00f9120b77494, topic:large-language-models, topic:llm
matched fp:a8c00f9120b77494, topic:instruction-tuning
matched fp:a8c00f9120b77494, topic:multimodal, name:multimodal, desc:multimodal