Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
[ICCV 2025] LLaVA-CoT, a visual language model capable of spontaneous, systematic reasoning
| Date | Stars |
|---|---|
| 2026-07-31 | 2134 |
| 2026-08-05 | 2134 |
| 2026-08-06 | 2134 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
<div align=center> <img src="figures/logo.png" width="280px"> </div> <h2 align="center"> <a href="https://arxiv.org/abs/2411.10440">LLaVA-CoT: Let Vision Language Models Reason Step-by-Step</a></h2> <h5 align="center"> If you like our project, please give us a star ⭐ on GitHub for the latest update.</h5> <h5 align=center> [](https://huggingface.co/Xkev/Llama-3.2V-11B-cot) [](https://arxiv.org/abs/2411.10440) [](https://x.com/Kevin_GuoweiXu/status/1858338565463421244) [](https://github.com/PKU-YuanGroup/LLaVA-CoT/blob/main/LICENSE) [](https://replicate.com/chenxwh/llava-cot) [](https://wisemodel.cn/codes/KevinTHU/LLaVA-CoT) </h5> <h5 align="center">本项目希望构建开源多模态慢思考推理大模型,由北大深研院袁粒老师课题组研发。</h5> ## 📣 News - **[2025/07/01]** LLaVA-CoT has been accepted by ICCV 2025! - **[2025/04/13]** We have updated the inference script that does not rely on VLMEvalKit, based on the suggestion in [this issue](https://github.com/PKU-YuanGroup/LLaVA-CoT/issues/34). - **[2025/01/08]** We released the full training code. - **[2025/01/02]** We discovered that when testing with the AI2D benchmark, we were using AI2D_TEST_NO_MASK, while the VLMEvalKit utilizes AI2D_TEST. We previously overlooked the distinction between the two, and we sincerely apologize for this oversight. We will make the necessary corrections. - **[2024/11/28]** We've released the dataset: [https://huggingface.co/datasets/Xkev/LLaVA-CoT-100k](<[dataset_generation/generate.py](https://huggingface.co/datasets/Xkev/LLaVA-CoT-100k)>) - **[2024/11/25]** We've released the code for dataset generation: [dataset_generation/generate.py](dataset_generation/generate.py) - **[2024/11/23]** We've released the Gradio App: [https://huggingface.co/spaces/Xkev/Llama-3.2V-11B-cot](https://huggingface.co/spaces/Xkev/Llama-3.2V-11B-cot) - **[2024/11/20]** We've released the pretrained weights: [https://huggingface.co/Xkev/Llama-3.2V-11B-cot](https://huggingface.co/Xkev/Llama-3.2V-11B-cot) - **[2024/11/18]** We've released our paper: [https://arxiv.org/abs/2411.10440](https://arxiv.org/abs/2411.10440) - **[2024/11/18]** Welcome to **watch** 👀 this repository for the latest updates. ## 🔥 Highlights **LLaVA-CoT** is a visual language model capable of spontaneous, systematic reasoning. Our 11B model outperforms **Gemini-1.5-pro**,**GPT-4o-mini**, and **Llama-3.2-90B-Vision-Instruct** on six challenging multimodal benchmarks. <div align=center> <img src="figures/result.png" width="300px"> </div> ## 🚀 Demos LLaVA-CoT begins by outlining the problem, interprets relevant information from the image, proceeds step-by-step through reasoning, and ultimately reaches a well-supported conclusion. ### Reasoning Problems | **Question** | <img src="figures/reasoning.png" width="400"> <br> Subtract all tiny shiny balls. Subtract all purple objects. How many objects are left? Options: A. 4, B. 8, C. 2, D. 6
Excerpt of 21,370 characters
Read on GitHub32
Chenxi
2
1
1
1
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:e482b45267674b3d, llm:Description: "LLaVA-CoT, a visual language model capable of spontaneous, systematic reasoning"; ICCV 2025 paper — visual language model (multimodal) focused on chain-of-thought reasoning.
matched fp:e482b45267674b3d, llm:Description: "LLaVA-CoT, a visual language model capable of spontaneous, systematic reasoning"; ICCV 2025 paper — visual language model (multimodal) focused on chain-of-thought reasoning.
matched fp:e482b45267674b3d, llm:Description: "LLaVA-CoT, a visual language model capable of spontaneous, systematic reasoning"; ICCV 2025 paper — visual language model (multimodal) focused on chain-of-thought reasoning.
matched fp:e482b45267674b3d, llm:Description: "LLaVA-CoT, a visual language model capable of spontaneous, systematic reasoning"; ICCV 2025 paper — visual language model (multimodal) focused on chain-of-thought reasoning.