Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
| Date | Stars |
|---|---|
| 2026-07-31 | 1305 |
| 2026-08-03 | 1305 |
| 2026-08-06 | 1305 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
<p align="center">
<img src="https://github.com/DAMO-NLP-SG/VideoLLaMA2/blob/e7bc34e0e9a96d77947a75b54399d9f96ccf209d/assets/logo.png" width="150" style="margin-bottom: 0.2;"/>
<p>
<h3 align="center"><a href="https://arxiv.org/abs/2406.07476" style="color:#9C276A">
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs</a></h3>
<h5 align="center"> If our project helps you, please give us a star ⭐ on GitHub to support us. 🙏🙏 </h2>
<h5 align="center">
[](https://huggingface.co/spaces/lixin4ever/VideoLLaMA2-AV)
[](https://huggingface.co/spaces/lixin4ever/VideoLLaMA2)
[](https://huggingface.co/collections/DAMO-NLP-SG/videollama-2-6669b6b6f0493188305c87ed)
[](https://huggingface.co/datasets/DAMO-NLP-SG/Multi-Source-Video-Captioning) <br>
[](https://github.com/DAMO-NLP-SG/VideoLLaMA2/blob/main/LICENSE)
[](https://hits.seeyoufarm.com)
[](https://github.com/DAMO-NLP-SG/VideoLLaMA2/issues?q=is%3Aopen+is%3Aissue)
[](https://github.com/DAMO-NLP-SG/VideoLLaMA2/issues?q=is%3Aissue+is%3Aclosed) <br>
[](https://huggingface.co/papers/2406.07476)
[](https://arxiv.org/abs/2406.07476) <br>
</h5>
[](https://paperswithcode.com/sota/zero-shot-video-question-answer-on-egoschema-1?p=videollama-2-advancing-spatial-temporal) <br>
[](https://paperswithcode.com/sota/video-question-answering-on-perception-test?p=videollama-2-advancing-spatial-temporal) <br>
[](https://paperswithcode.com/sota/video-question-answering-on-mvbench?p=videollama-2-advancing-spatial-temporal) <br>
[](https://paperswithcode.com/sota/zero-shot-video-question-answer-on-video-mme-1?p=videollama-2-advancing-spatial-temporal) <br>
[](https://paperswithcode.com/sota/zero-shot-video-question-answer-on-video-mme?p=videollama-2-advancing-spatial-temporal) <br>
<details open><summary>💡 Some other multimodal-LLM projects from our team may interest you ✨. </summary><p>
<!-- may -->
> [**VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding**](https://github.com/DAMO-NLP-SG/VideoLLaMA3) <br>
> Boqiang Zhang<sup>* </sup>, Kehan Li<sup>* </sup>, Zesen Cheng<sup>* </sup>, Zhiqiang Hu<sup>* </sup>, Yuqian Yuan<sup>* </sup>, Guanzheng Chen<sup>* </sup>, Sicong Leng<sup>* </sup>, Yuming Jiang<sup>* </sup>, Hang Zhang<sup>* </sExcerpt of 26,314 characters
Read on GitHub98
LI XIN · @alibaba
25
10
3
2
2
1
1
lionHC
1
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:4a5fbbcf05b37936, llm:Repository title and description: 'VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs' — focuses on video and audio understanding in large language models for video.
matched fp:4a5fbbcf05b37936, llm:Repository title and description: 'VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs' — focuses on video and audio understanding in large language models for video.
matched fp:4a5fbbcf05b37936, llm:Repository title and description: 'VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs' — focuses on video and audio understanding in large language models for video.
matched fp:4a5fbbcf05b37936, llm:Repository title and description: 'VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs' — focuses on video and audio understanding in large language models for video.