Top AI Repos โ open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
[ACL 2025 ๐ฅ] Rethinking Step-by-step Visual Reasoning in LLMs
| Date | Stars |
|---|---|
| 2026-07-31 | 307 |
| 2026-08-06 | 307 |
Today
โ stars today
This week
โ stars this week
This month
โ stars this month
Momentum
0.0
growth rate 0.00%/day
<div align=center> <img src="figures/logo2.png" width="100px"> </div> <h2 align="center"> LlamaV-o1: Rethinking Step-By-Step Visual Reasoning in LLMs [ACL 2025 ๐ฅ]</h2>  [Omkar Thawakar](https://omkarthawakar.github.io/)* , [Dinura Dissanayake](https://github.com/mbzuai-oryx/LlamaV-o1)* , [Ketan More](https://github.com/mbzuai-oryx/LlamaV-o1)* , [Ritesh Thawkar](https://github.com/mbzuai-oryx/LlamaV-o1)* , [Ahmed Heakl](https://scholar.google.com/citations?user=JcWO9OUAAAAJ&hl=en)* , [Noor Ahsan](https://github.com/mbzuai-oryx/LlamaV-o1)* , [Yuhao Li](https://github.com/mbzuai-oryx/LlamaV-o1)* , [Mohammed Zumri](https://github.com/mbzuai-oryx/LlamaV-o1)* , [Jean Lahoud](https://scholar.google.com/citations?user=LsivLPoAAAAJ&hl=en)*, [Rao Muhammad Anwer](https://scholar.google.com/citations?hl=en&authuser=1&user=_KlvMVoAAAAJ), [Hisham Cholakkal](https://scholar.google.com/citations?hl=en&user=bZ3YBRcAAAAJ), [Ivan Laptev](https://mbzuai.ac.ae/study/faculty/ivan-laptev/), [Mubarak Shah](https://www.cs.ucf.edu/person/mubarak-shah/), [Fahad Shahbaz Khan](https://scholar.google.es/citations?user=zvaeYnUAAAAJ&hl=en) and [Salman Khan](https://salman-h-khan.github.io/) *Equal Contribution **Mohamed bin Zayed University of Artificial Intelligence, UAE** <h5 align="center"> If you like our project, please give us a star โญ on GitHub for the latest update.</h5> ## ๐ฃ Latest Updates - **May-15-2025**: LlamaV-o1 Accepted in ACL-2025 finding. - **January-13-2025**: Technical Report of LlamaV-o1 is released on Arxiv. [Paper](https://arxiv.org/abs/2501.06186) - **January-10-2025**: *Code, Model & Dataset release. Our VCR-Bench is available at: [HuggingFace](https://huggingface.co/datasets/omkarthawakar/VRC-Bench). Model Checkpoint: [HuggingFace](https://huggingface.co/omkarthawakar/LlamaV-o1). Code is available at: [GitHub](https://github.com/mbzuai-oryx/LlamaV-o1/).๐ค --- ## ๐ฅ Highlights **LlamaV-o1** is a Large Multimodal Model capable of spontaneous reasoning. - Our LlamaV-o1 model outperforms **Gemini-1.5-flash**,**GPT-4o-mini**, **Llama-3.2-Vision-Instruct**, **Mulberry**, and **Llava-CoT** on our proposed VCR-Bench. - Our LlamaV-o1 model outperforms **Gemini-1.5-Pro**,**GPT-4o-mini**, **Llama-3.2-Vision-Instruct**, **Mulberry**, **Llava-CoT**, etc. on six challenging multimodal benchmarks (MMStar, MMBench, MMVet, MathVista, AI2D and Hallusion). ## Contributions ๐ - Step-by-Step Visual Reasoning Benchmark: To the best of our knowledge, the proposed benchmark is the first effort designed to evaluate multimodal multi-step reasoning tasks across diverse topics. The proposed benchmark, named VRC-Bench, spans around eight different categories (Visual Reasoning, Math & Logic Reasoning, Social & Cultural Context, Medical Imaging (Basic Medical Science), Charts & Diagram Understanding, OCR & Document Understanding, Complex Visual Perception and Scientific Reasoning) with over 1,000 challenging samples and more than 4k reasoning steps. - Novel Evaluation Metric: A metric that assesses the reasoning quality at the level of individual steps, emphasizing both correctness and logical coherence. - Combined Multi-Step Curriculum Learning and Beam Search Approach: A multimodal rea- soning method, named LlamaV-o1, that combines the structured progression of curriculum learning with the efficiency of Beam Search. The proposed approach ensures incremental skill development while optimizing reasoning paths, enabling the model to be effective in complex multi-step visual reasoning tasks in terms of both accuracy and efficiency. Specifi- cally, the proposed LlamaV-o1 achieves an absolute gain of 3.8% in terms of average score across six benchmarks while being 5ร faster, compared to the recent Llava-CoT. --- ### Dataset Overview <div align=center> <img src="figures/dataset_overview.png" width="900px"> </div> The figure presents our benchmark structure and the comparative performance of LMMs on VRC-Bench. The
Excerpt of 11,829 characters
Read on GitHubWould you bet a product on this? Bounded 0โ100 and slow moving.
matched fp:f57dc7cc5d1a9b36, llm:Repository title and description: 'LlamaV-o1' and '[ACL 2025] Rethinking Step-by-step Visual Reasoning in LLMs' โ indicates work on visual reasoning with large language models (multimodal), likely research code for vision+LLM step-by-step reasoning.
matched fp:f57dc7cc5d1a9b36, llm:Repository title and description: 'LlamaV-o1' and '[ACL 2025] Rethinking Step-by-step Visual Reasoning in LLMs' โ indicates work on visual reasoning with large language models (multimodal), likely research code for vision+LLM step-by-step reasoning.
matched fp:f57dc7cc5d1a9b36, llm:Repository title and description: 'LlamaV-o1' and '[ACL 2025] Rethinking Step-by-step Visual Reasoning in LLMs' โ indicates work on visual reasoning with large language models (multimodal), likely research code for vision+LLM step-by-step reasoning.
matched fp:f57dc7cc5d1a9b36, llm:Repository title and description: 'LlamaV-o1' and '[ACL 2025] Rethinking Step-by-step Visual Reasoning in LLMs' โ indicates work on visual reasoning with large language models (multimodal), likely research code for vision+LLM step-by-step reasoning.