Top AI Repos โ open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
๐๏ธ + ๐ฌ + ๐ง = ๐ค Curated list of top foundation and multimodal models! [Paper + Code + Examples + Tutorials]
| Date | Stars |
|---|---|
| 2026-07-24 | 638 |
| 2026-07-25 | 638 |
| 2026-07-28 | 638 |
| 2026-07-30 | 638 |
| 2026-08-06 | 638 |
Today
โ stars today
This week
โ stars this week
This month
โ stars this month
Momentum
0.0
growth rate 0.00%/day
<h1 align="center">awesome foundation and multimodal models</h1>
## ๐๏ธ + ๐ฌ + ๐ง = ๐ค
**foundation model** - a pre-trained machine learning model that serves as a base for a wide range of downstream tasks. It captures general knowledge from a large dataset and can be fine-tuned to perform specific tasks more effectively.
**multimodal model** - a model that can process multiple modalities (e.g. text, image,
video, audio, etc.) at the same time.
## ๐ค models
<!--- AUTOGENERATED_PAPERS_LIST -->
<!---
WARNING: DO NOT EDIT THIS LIST MANUALLY. IT IS AUTOMATICALLY GENERATED.
HEAD OVER TO CONTRIBUTING.MD FOR MORE DETAILS ON HOW TO MAKE CHANGES PROPERLY.
-->
### YOLO-World: Real-Time Open-Vocabulary Object Detection
[](https://arxiv.org/abs/2401.17270) [](https://github.com/AILab-CVC/YOLO-World) [](https://youtu.be/X7gKBGVz4vs) [](https://huggingface.co/spaces/SkalskiP/YOLO-World) [](https://colab.research.google.com/github/roboflow-ai/notebooks/blob/main/notebooks/zero-shot-object-detection-with-yolo-world.ipynb)
Tianheng Cheng, Lin Song, Yixiao Ge, Wenyu Liu, Xinggang Wang, Ying Shan
- **Date:** 2024-01-30
- **Modalities:** ๐๏ธ + ๐ฌ
- **Tasks:** Zero-Shot Object Detection
### Depth Anything
[](https://arxiv.org/abs/2401.10891) [](https://github.com/LiheYoung/Depth-Anything) [](https://huggingface.co/spaces/LiheYoung/Depth-Anything) [](https://huggingface.co/LiheYoung/depth_anything_vitl14) [](https://colab.research.google.com/github/NielsRogge/Transformers-Tutorials/blob/master/Depth%20Anything/Predicting_depth_in_an_image_with_Depth_Anything.ipynb)
Lihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu, Jiashi Feng, Hengshuang Zhao
- **Date:** 2024-01-19
- **Modalities:** ๐
- **Tasks:** Depth Estimation
### EfficientSAM: Leveraged Masked Image Pretraining for Efficient Segment Anything
[](https://arxiv.org/abs/2312.00863) [](https://github.com/yformer/EfficientSAM) [](https://huggingface.co/spaces/SkalskiP/EfficientSAM) [](https://huggingface.co/merve/EfficientSAM)
Yunyang Xiong, Bala Varadarajan, Lemeng Wu, Xiaoyu Xiang, Fanyi Xiao, Chenchen Zhu, Xiaoliang Dai, Dilin Wang, Fei Sun, Forrest Iandola, Raghuraman Krishnamoorthi, Vikas Chandra
- **Date:** 2023-12-01
- **Modalities:** ๐๏ธ
- **Tasks:** Zero-Shot Object Segmentation
### Qwen-VL-Plus / Max
[](https://arxiv.org/abs/2308.12966) [](https://github.com/QwenLM/Qwen-VL#qwen-vl-plus) [](https://huggingface.co/spaces/Qwen/Qwen-VL-Plus) [](https://huggingface.co/Qwen/Qwen-VL)
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, Jingren Zhou
- **Date:** 2023-11-28
- **Modalities:** ๐๏ธ + ๐ฌ
- **Tasks:** Image Captioning, VQA, ZerExcerpt of 20,738 characters
Read on GitHubPiotr Skalski ยท @roboflow
19
Merve Noyan ยท @huggingface ยท France
11
1
Would you bet a product on this? Bounded 0โ100 and slow moving.
matched fp:86fd1e805c1f01d1, topic:multimodal, topic:clip, topic:image-captioning