Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
[ECCV2024] Video Foundation Models & Data for Multimodal Understanding
| Date | Stars |
|---|---|
| 2026-07-24 | 2342 |
| 2026-07-25 | 2342 |
| 2026-07-28 | 2342 |
| 2026-07-30 | 2342 |
| 2026-07-31 | 2343 |
| 2026-08-06 | 2343 |
Today
— stars today
This week
+1 stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.04%/day
# InternVideo: Video Foundation Models for Multimodal Understanding </div> --- <div align='center'> <img src="./InternVideo2/figs/teaser-internvideo2.png" class="interpolation-image" alt="internvideo2_performance." height="96%" width="96%" /> </div> This repo contains InternVideo series and related works in video foundation models. - [InternVideo](InternVideo1): general video foundation models via generative and discriminative learning - [InternVideo2](InternVideo2): scaling video foundation models for multimodal video understanding - [InternVideo2.5](InternVideo2.5): empowering video mllms with long and rich context modeling - [InternVideo3](InternVideo3): multimodal contextual reasoning via efficient long-horizon agents - [InternVideo-Next](InternVideo-Next): general video foundation models for genuine world understanding - [InternVid](Data/InternVid): a large-scale video-text dataset for multimodal understanding and generation ## Updates - `2026.06`: [InternVideo3](InternVideo3) is released with the [technical report](https://arxiv.org/abs/2606.12195), [8B instruct model](https://huggingface.co/yanziang/InternVideo3-8B-Instruct), [long-video SFT dataset](https://huggingface.co/datasets/yanziang/InternVideo3_Dataset), evaluation scripts, and an initial video-agent implementation in [Vidify](https://github.com/shepnerd/vidify). - `2025.12`: [InternVideo-Next](InternVideo-Next) is released with the [technical report](https://arxiv.org/abs/2512.01342), pretrained model weights in the [Hugging Face collection](https://huggingface.co/collections/OpenGVLab/internvideo-next), and pretraining code. - `2025.01`: [InternVideo2.5](InternVideo2.5) is now released! Check out the [technical report](https://arxiv.org/pdf/2501.12386) for detailed insights, and access the [model](https://huggingface.co/OpenGVLab/InternVL_2_5_HiCo_R16) on HuggingFace. - `2024.08.12`: We provide smaller models, [InternVideo2-S/B/L](./InternVideo2/single_modality/MODEL_ZOO.md), which are distilled from InternVideo2-1B. We also build smaller [VideoCLIP](./InternVideo2/multi_modality/MODEL_ZOO.md) with MobileCLIP. - `2024.08`: [InternVideo2-Stage3-8B](https://huggingface.co/OpenGVLab/InternVideo2-Chat-8B) and [InternVideo2-Stage3-8B-HD](https://huggingface.co/OpenGVLab/InternVideo2_chat_8B_HD) are released. 8B indicates the use of InternVideo2-1B and the 7B LLM. - `2024.07`: The video annotation for InternVid2 ([HuggingFace](https://huggingface.co/datasets/OpenGVLab/InternVideo2_Vid_Text)) is released. - `2024.06`: The full version of the video annotation (230M video-text pairs) for InternVid ([OpenDataLab](https://opendatalab.com/shepshep/InternVidFull) | [HuggingFace](https://huggingface.co/datasets/OpenGVLab/InternVid-Full)) is released. - `2024.04`: The [Checkpoints](https://huggingface.co/collections/OpenGVLab/internvideo2-6618ccb574bd2f91410df5cd) and scripts for InternVideo2 are released. - `2024.03`: The [technical report](https://arxiv.org/abs/2403.15377) of InternVideo2 is released. - `2024.01`: [InternVid](Data/InternVid) (a video-text dataset for video understanding and generation) has been accepted for spotlight presentation of ICLR 2024. - `2023.07`: A **video-text dataset InternVid** is released at [here](Data/InternVid) for facilitating multimodal understanding and generation. - `2023.05`: **Video instruction data** are released at [here](Data/instruction_data) for tuning end-to-end video-centric multimodal dialogue systems like [VideoChat](https://github.com/OpenGVLab/Ask-Anything). - `2023.01`: The [code & models](InternVideo1) of InternVideo are released. - `2022.12`: The [technical report](https://arxiv.org/pdf/2212.03191.pdf) of InternVideo is released. - `2022.09`: Press releases of InternVideo ([official](https://www.shlab.org.cn/news/5443279) | [163 news](https://www.163.com/dy/article/HG939TNR0530QRMB.html) | [qq news](https://new.qq.com/rain/a/20220902A053JP00)). ## Contact - If you have any questions during the trial, running or
Excerpt of 4,592 characters
Read on GitHub125
Yinan He · @OpenGVLab · China
44
29
17
9
9
6
4
Bingkun Huang · Google
3
2
2
1
Masoud Kaviani · Sharif University of Technology · Iran
1
Rong Ou · @NVIDIA · United States
1
vansin · United States
1
1
1
1
1
1
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:7bea1cdd3ae51f72, topic:multimodal, desc:multimodal, readme:multimodal
matched fp:7bea1cdd3ae51f72, topic:instruction-tuning
matched fp:7bea1cdd3ae51f72, topic:foundation-models, readme:model weights, readme:pretrained model
matched fp:7bea1cdd3ae51f72, topic:contrastive-learning