Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
A family of lightweight multimodal models.
| Date | Stars |
|---|---|
| 2026-07-24 | 1053 |
| 2026-07-25 | 1053 |
| 2026-07-28 | 1053 |
| 2026-07-30 | 1053 |
| 2026-08-06 | 1053 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# Bunny: A family of lightweight multimodal models <p align="center"> <img src="./icon.png" alt="Logo" width="350"> </p> 📖 [Technical report](https://arxiv.org/abs/2402.11530) | 🤗 [Data](https://huggingface.co/datasets/BoyaWu10/Bunny-v1_1-data) | 🤖 [Data](https://www.modelscope.cn/datasets/BoyaWu10/Bunny-v1.1-data) | 🤗 [HFSpace](https://huggingface.co/spaces/BoZhaoHuggingFace/Bunny) 🐰 [Demo](http://bunny.baai.ac.cn) **Bunny-Llama-3-8B-V**: 🤗 [v1.1](https://huggingface.co/BAAI/Bunny-v1_1-Llama-3-8B-V) | 🤗 [v1.0](https://huggingface.co/BAAI/Bunny-Llama-3-8B-V) | 🤗 [v1.0-GGUF](https://huggingface.co/BAAI/Bunny-Llama-3-8B-V-gguf) **Bunny-4B**: 🤗 [v1.1](https://huggingface.co/BAAI/Bunny-v1_1-4B) | 🤗 [v1.0](https://huggingface.co/BAAI/Bunny-v1_0-4B) | 🤗 [v1.0-GGUF](https://huggingface.co/BAAI/Bunny-v1_0-4B-gguf) Bunny is a family of lightweight but powerful multimodal models. It offers multiple plug-and-play vision encoders, like **EVA-CLIP, SigLIP** and language backbones, including **Llama-3-8B, Phi-3-mini, Phi-1.5, StableLM-2, Qwen1.5, MiniCPM and Phi-2**. To compensate for the decrease in model size, we construct more informative training data by curated selection from a broader data source. We are thrilled to introduce **Bunny-Llama-3-8B-V**, the pioneering vision-language model based on Llama-3, showcasing exceptional performance. The v1.1 version accepts high-resolution images up to **1152x1152**.  Moreover, our **Bunny-4B** model built upon SigLIP and Phi-3-mini outperforms the state-of-the-art MLLMs, not only in comparison with models of similar size but also against larger MLLMs (7B and 13B). Also, the v1.1 version accepts high-resolution images up to **1152x1152**. <details> <summary>Expand to see the performance of Bunny-4B</summary> <IMG src="comparison_4B.png"/> </details> ## News and Updates * 2024.07.23 🔥 **All of the training strategy and data of latest Bunny is released!** Check more details about Bunny in [Technical Report](https://arxiv.org/abs/2402.11530), [Data](https://huggingface.co/datasets/BoyaWu10/Bunny-v1_1-data) and [Training Tutorial](#training-tutorial)! * 2024.07.21 🔥 **SpatialBot, SpatialQA and SpatialBench are released!** SpatialBot is an embodiment model based on Bunny, which comprehends spatial relationships by understanding and using depth information. Try model, dataset and benchmark at [GitHub](https://github.com/BAAI-DCAI/SpatialBot)! * 2024.06.20 🔥 **MMR benchmark is released!** It is a benchmark for measuring MLLMs' understanding ability and their robustness against misleading questions. Check the performance of Bunny and more details in [GitHub](https://github.com/BAAI-DCAI/Multimodal-Robustness-Benchmark)! * 2024.06.01 🔥 **Bunny-v1.1-Llama-3-8B-V, supporting 1152x1152 resolution, is released!** It is built upon SigLIP and Llama-3-8B-Instruct with S$`^2`$-Wrapper. Check more details in [HuggingFace](https://huggingface.co/BAAI/Bunny-v1_1-Llama-3-8B-V) and [wisemodel](https://wisemodel.cn/models/BAAI/Bunny-v1.1-Llama-3-8B-V)! 🐰 [Demo](http://bunny.baai.ac.cn) * 2024.05.08 **Bunny-v1.1-4B, supporting 1152x1152 resolution, is released!** It is built upon SigLIP and Phi-3-Mini-4K 3.8B with S$`^2`$-Wrapper. Check more details in [HuggingFace](https://huggingface.co/BAAI/Bunny-v1_1-4B)! 🐰 [Demo](http://bunny.baai.ac.cn) * 2024.05.01 **Bunny-v1.0-4B, a vision-language model based on Phi-3, is released!** It is built upon SigLIP and Phi-3-Mini-4K 3.8B. Check more details in [HuggingFace](https://huggingface.co/BAAI/Bunny-v1_0-4B)! 🤗 [GGUF](https://huggingface.co/BAAI/Bunny-v1_0-4B-gguf) * 2024.04.21 **Bunny-Llama-3-8B-V, the first vision-language model based on Llama-3, is released!** It is built upon SigLIP and Llama-3-8B-Instruct. Check more details in [HuggingFace](https://huggingface.co/BAAI/Bunny-Llama-3-8B-V), [ModelScope](https://www.modelscope.cn/models/BAAI/Bunny-Llama-3-8B-V), and [wisemodel](https://wisemodel.cn/models/BAAI/Bunny-Llam
Excerpt of 35,666 characters
Read on GitHub82
19
5
4
2
ambivalent · Vietnam
1
1
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:04a82bd77dc384d6, topic:vlm, desc:multimodal, readme:multimodal
matched fp:04a82bd77dc384d6, topic:chatgpt