Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
[ICLR'24] Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
| Date | Stars |
|---|---|
| 2026-07-31 | 297 |
| 2026-08-02 | 297 |
| 2026-08-06 | 297 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning [[ICLR 2024](https://iclr.cc)] [Fuxiao Liu](https://fuxiaoliu.github.io/), [Kevin Lin](https://sites.google.com/site/kevinlin311tw/me), [Linjie Li](https://www.microsoft.com/en-us/research/people/linjli/), [Jianfeng Wang](http://jianfengwang.me/), [Yaser Yacoob](https://www.umiacs.umd.edu/people/yaser), [Lijuan Wang](https://www.microsoft.com/en-us/research/people/lijuanw/) [[Project Page](https://fuxiaoliu.github.io/LRV/)] [[Paper](http://arxiv.org/abs/2306.14565)] You can compare between our models and original models below. If the online demos don't work, please email `[email protected]`. If you find our work interesting, please cite our work. Thanks!!! ```bibtex @article{liu2023aligning, title={Aligning Large Multi-Modal Model with Robust Instruction Tuning}, author={Liu, Fuxiao and Lin, Kevin and Li, Linjie and Wang, Jianfeng and Yacoob, Yaser and Wang, Lijuan}, journal={arXiv preprint arXiv:2306.14565}, year={2023} } @article{liu2023hallusionbench, title={HallusionBench: You See What You Think? Or You Think What You See? An Image-Context Reasoning Benchmark Challenging for GPT-4V (ision), LLaVA-1.5, and Other Multi-modality Models}, author={Liu, Fuxiao and Guan, Tianrui and Li, Zongxia and Chen, Lichang and Yacoob, Yaser and Manocha, Dinesh and Zhou, Tianyi}, journal={arXiv preprint arXiv:2310.14566}, year={2023} } @article{liu2023mmc, title={MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning}, author={Liu, Fuxiao and Wang, Xiaoyang and Yao, Wenlin and Chen, Jianshu and Song, Kaiqiang and Cho, Sangwoo and Yacoob, Yaser and Yu, Dong}, journal={arXiv preprint arXiv:2311.10774}, year={2023} } ``` ### Both LRV-V1 and LRV-V2 support training on V100 32GB. 📺 [[LRV-V2(Mplug-Owl) Demo](https://edfab153e1ff6d3c51.gradio.live)], [[mplug-owl Demo](https://huggingface.co/spaces/MAGAer13/mPLUG-Owl)] <br> 📺 [[LRV-V1(MiniGPT4) Demo](https://d225baa9dda7ba3877.gradio.live)], [[MiniGPT4-7B Demo](https://a7adeb59efb6b836f2.gradio.live)] ## Updates - [03/13]🔥 Our paper ["MMC: Advancing Multimodal Chart Understanding with LLM Instruction Tuning"](https://arxiv.org/pdf/2311.10774.pdf) is accepted to **[NAACL 2024](https://2024.naacl.org)**. - [02/26]🔥 Our paper ["HallusionBench: You See What You Think? Or You Think What You See? An Image-Context Reasoning Benchmark Challenging for GPT-4V(ision), LLaVA-1.5, and Other Multi-modality Models"](https://arxiv.org/abs/2310.14566) is accpeted to **[CVPR 2024](https://cvpr.thecvf.com)**. - [01/15]🔥 Our paper is accepted by **[ICLR 2024](https://iclr.cc)**. Camera-Ready Version will be ready soon! - [11/15]🔥 Our paper ["MMC: Advancing Multimodal Chart Understanding with LLM Instruction Tuning"](https://arxiv.org/pdf/2311.10774.pdf) is now available on Arxiv. - [10/24]🔥 Please check our new work to benchmark the **failure cases of GPT4V** ["HallusionBench: You See What You Think? Or You Think What You See? An Image-Context Reasoning Benchmark Challenging for GPT-4V(ision), LLaVA-1.5, and Other Multi-modality Models"](https://arxiv.org/abs/2310.14566)([repo](https://github.com/tianyi-lab/HallusionBench)). - [9/20] 🔥 More knowledge manipulation data will be release soon! - [8/24] 🔥 We release some visual instruction data (with knowledge manipulations) for chart images to increase the diversity of our dataset. [data](download.txt#L33) and [image](download.txt#L36). - [8/17] 🔥 Model weight of **LRV-Instruction V2** is available from [here](download.txt#L24). - [8/16] 🔥 We release additional **180k** visual instruction tuning data by generated GPT4. You can download from [here](download.txt#L20). Our LRV-Instruction dataset contains **320k** visual instruction data from in total. - [8/14] 🔥 We **manually clean** the dataset. The new version can be downloaded from [Training Set](download.txt#L5) and [Evaluation Set](Evaluation/evaluation_set.json).
Excerpt of 14,948 characters
Read on GitHubWould you bet a product on this? Bounded 0–100 and slow moving.
matched fp:a53c0a67806869e6, topic:foundation-models, topic:gpt, topic:llama