Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Collection of AWESOME vision-language models for vision tasks
| Date | Stars |
|---|---|
| 2026-07-24 | 3128 |
| 2026-07-25 | 3128 |
| 2026-07-28 | 3130 |
| 2026-07-30 | 3130 |
| 2026-07-31 | 3128 |
| 2026-08-06 | 3128 |
Today
— stars today
This week
-2 stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
## Awesome Vision-Language Models [](https://awesome.re) <img src="./images/overview.png" width="96%" height="96%"> This is the repository of **Vision Language Models for Vision Tasks: a Survey**, a systematic survey of VLM studies in various visual recognition tasks including image classification, object detection, semantic segmentation, etc. For details, please refer to: **Vision-Language Models for Vision Tasks: A Survey** [[Paper](https://arxiv.org/abs/2304.00685)] *IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2024* 🤩 Our paper is selected into **TPAMI Top 50 Popular Paper List** !! [](https://arxiv.org/abs/2304.00685) [](https://GitHub.com/Naereen/StrapDown.js/graphs/commit-activity) [](http://makeapullrequest.com) <!-- [](http://commonmark.org) --> <!-- [](http://ansicolortags.readthedocs.io/?badge=latest) --> *Feel free to pull requests or contact us if you find any related papers that are not included here.* The process to submit a pull request is as follows: - a. Fork the project into your own repository. - b. Add the Title, Paper link, Conference, Project/Code link in `README.md` using the following format: ``` |[Title](Paper Link)|Conference|[Code/Project](Code/Project link)| ``` - c. Submit the pull request to this branch. **We plan to update the arXiv version of our survey paper soon. If your paper is missing from this repository, feel free to contact us or open an issue!** ## 🔥 News <details open><summary>📣 We also have a collection on agentic MLLMs that may interest you ✨. </summary><p> > [**Awesome-Agentic-MLLMs**](https://github.com/HJYao00/Awesome-Agentic-MLLMs) <br> > [](https://arxiv.org/abs/2510.10991) [](https://github.com/HJYao00/Awesome-Agentic-MLLMs) </p ></details> 📅 Last update on 2025/10/14 #### VLMs and Synthetic Data * [CVPR 2025] Synthetic Data is an Elegant GIFT for Continual Vision-Language Models [[Paper](https://arxiv.org/pdf/2503.04229v1)][[Code](https://github.com/Luo-Jiaming/GIFT_CL)] * [CVPR 2025] Enhancing Vision-Language Compositional Understanding with Multimodal Synthetic Data [[Paper](https://openaccess.thecvf.com/content/CVPR2025/papers/Li_Enhancing_Vision-Language_Compositional_Understanding_with_Multimodal_Synthetic_Data_CVPR_2025_paper.pdf)] * [OpenReview 2025] A Survey on Bridging VLMs and Synthetic Data [[Paper](https://openreview.net/pdf?id=ThjDCZOljE)][[Code](https://github.com/mghiasvand1/Awesome-VLM-Synthetic-Data)] #### VLM Pre-training Methods * [NeurIPS 2024] PLIP: Language-Image Pre-training for Person Representation Learning [[Paper](https://papers.nips.cc/paper_files/paper/2024/file/510ad3018bbdc5b6e3b10646e2e35771-Paper-Conference.pdf)][[Code](https://github.com/Zplusdragon/PLIP)] * [NeurIPS 2024] LoTLIP: Improving Language-Image Pre-training for Long Text Understanding [[Paper](https://papers.nips.cc/paper_files/paper/2024/file/77828623211df05497ce3658300dafd9-Paper-Conference.pdf)][[Code](https://wuw2019.github.io/lot-lip)] * [NeurIPS 2024] Accelerating Pre-training of Multimodal LLMs via Chain-of-Sight [[Paper](https://papers.nips.cc/paper_files/paper/2024/file/8a54a80ffc2834689ffdd0920202018e-Paper-Conference.pdf)][[Code](https://chain-of-sight.github.io/)] #### VLM Transfer Learning Methods * [ICCV 2025] One Last Attention for Your Vision-Language Model [[Paper](https://arxiv.org/pdf/2507.15480v1)][[Code](https://github.com/khufia/RAda/tree/main)] *
Excerpt of 59,024 characters
Read on GitHub92
Peng "Richard" Xia · UNC-Chapel Hill · United States
5
Mohammad Ghiasvand Mohammadkhani
3
Masoud Jafaripour · CS @ UofA · Canada
2
Jaehyun Jang · KAIST(Korea Advanced Institute of Science and Technology) · South Korea
1
1
1
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:7776592a8f9af524, topic:vision-language-model, topic:clip, readme:multimodal
matched fp:7776592a8f9af524, topic:computer-vision, readme:object detection, readme:semantic segmentation