Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
| Date | Stars |
|---|---|
| 2026-07-31 | 549 |
| 2026-08-01 | 549 |
| 2026-08-06 | 549 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
[](https://GitHub.com/Naereen/StrapDown.js/graphs/commit-activity)
[](http://makeapullrequest.com)
[](https://github.com/sindresorhus/awesome)
# Foundational Models Defining a New Era in Vision: A Survey and Outlook
Accepted for publication by **TPAMI** (IEEE Transactions on Pattern Analysis and Machine Intelligence).
> [**Foundational Models Defining a New Era in Vision: A Survey and Outlook**](https://arxiv.org/abs/2307.13721)<br>
> [Muhammad Awais](awaisrauf.github.io), [Muzammal Naseer](https://muzammal-naseer.netlify.app), [Salman Khan](https://salman-h-khan.github.io), [Rao Muhammad Anwer](https://scholar.google.fi/citations?user=_KlvMVoAAAAJ&hl=en), [Hisham Cholakkal](https://scholar.google.com/citations?user=bZ3YBRcAAAAJ&hl=en), [Mubarak Shah](https://www.crcv.ucf.edu/person/mubarak-shah/), [ Ming-Hsuan Yang](http://faculty.ucmerced.edu/mhyang/), [Fahad Shahbaz Khan](https://sites.google.com/view/fahadkhans/home)
> **<p align="justify"> Abstract:** *Vision systems to see and reason about the compositional nature of visual scenes are fundamental to understanding our
world. The complex relations between objects and their locations, ambiguities, and variations in the real-world environment can be
better described in human language, naturally governed by grammatical rules and other modalities such as audio and depth. The
models learned to bridge the gap between such modalities coupled with large-scale training data facilitate contextual reasoning,
generalization, and prompt capabilities at test time. These models are referred to as foundational models. The output of such models
can be modified through human-provided prompts without retraining, e.g., segmenting a particular object by providing a bounding box,
having interactive dialogues by asking questions about an image or video scene or manipulating the robot’s behavior through language
instructions. In this survey, we provide a comprehensive review of such emerging foundational models, including typical architecture
designs to combine different modalities (vision, text, audio, etc), training objectives (contrastive, generative), pre-training datasets,
fine-tuning mechanisms, and the common prompting patterns; textual, visual, and heterogeneous. We discuss the open challenges
and research directions for foundational models in computer vision, including difficulties in their evaluations and benchmarking, gaps in
their real-world understanding, limitations of their contextual understanding, biases, vulnerability to adversarial attacks, and
interpretability issues. We review recent developments in this field, covering a wide range of applications of foundation models
systematically and comprehensively.* </p>
<div align='center'>
<img src="overview.svg" width="60%" height="60%">
</div>
# <p align=center>`Awesome CV Foundational Models`</p>
A curated list of awesome foundational models in computer vision. This repo supplements our survey paper. We intend to continuously update it.
We strongly encourage authors of relevant works to make a pull request and add their paper's information.
____
## Citation
If you find our work useful in your research, please consider citing:
```
@article{awais2023foundational,
title={Foundational Models Defining a New Era in Vision: A Survey and Outlook},
author={Awais, Muhammad and Naseer, Muzammal and Khan, Salman and Anwer, Rao Muhammad and Cholakkal, Hisham and Shah, Mubarak and Yang, Ming-Hsuan and Khan, Fahad Shahbaz},
journal={arXiv preprint arXiv:2307.13721},
year={2023}
}
```
# Menu
- [Surveys](#surveys)
- [Year 2021](#2021)
- [Year 2022](#2022)
- [Year 2023](#2023)
# Surveys
[**Foundational Models Defining a New Era in Vision: A Survey and Outlook**](Excerpt of 192,782 characters
Read on GitHub21
9
3
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:73e168452c0ffa9e, llm:Repository name: 'Awesome-CV-Foundational-Models' (implies curated list of CV foundational models). No topics or README provided.
matched fp:73e168452c0ffa9e, llm:Repository name: 'Awesome-CV-Foundational-Models' (implies curated list of CV foundational models). No topics or README provided.