Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
A curated list of deep learning resources for video-text retrieval.
| Date | Stars |
|---|---|
| 2026-07-31 | 644 |
| 2026-08-02 | 644 |
| 2026-08-06 | 644 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# Awesome Video-Text Retrieval by Deep Learning [](https://github.com/sindresorhus/awesome) A curated list of deep learning resources for video-text retrieval. ## Contributing Please feel free to [pull requests](https://github.com/danieljf24/awesome-video-text-retrieval/pulls) to add papers. Markdown format: ```markdown - `[Author Journal/Booktitle Year]` Title. Journal/Booktitle, Year. [[paper]](link) [[code]](link) [[homepage]](link) ``` ## Table of Contents - [Implementations](#implementations) - [PyTorch](#pytorch) - [TensorFlow](#tensorflow) - [Others](#others) - [Papers](#papers) - [2023](#2023) - [2022](#2022) - [2021](#2021) - [2020](#2020) - [2019](#2019) - [2018](#2018) - [Before](#before) - [Ad-hoc Video Search](#ad-hoc-video-search) - [Other Related](#other-related) - [Datasets](#datasets) ## Implementations #### PyTorch - [hybrid_space](https://github.com/danieljf24/hybrid_space) - [dual_encoding](https://github.com/danieljf24/dual_encoding) - [w2vvpp](https://github.com/li-xirong/w2vvpp) - [Mixture-of-Embedding-Experts](https://github.com/antoine77340/Mixture-of-Embedding-Experts) - [howto100m](https://github.com/antoine77340/howto100m) - [collaborative](https://github.com/albanie/collaborative-experts) - [hgr](https://github.com/cshizhe/hgr_v2t) - [coot](https://github.com/gingsi/coot-videotext) - [mmt](https://github.com/gabeur/mmt) - [ClipBERT](https://github.com/jayleicn/ClipBERT) #### TensorFlow - [jsfusion](https://github.com/yj-yu/lsmdc) #### Others - [w2vv](https://github.com/danieljf24/w2vv)(Keras) #### Useful Toolkit - [Extracting CNN features from video frames by MXNet](https://github.com/xuchaoxi/video-cnn-feat) ## Papers ### 2023 - `[Pei et al. CVPR23]` CLIPPING: Distilling CLIP-Based Models with a Student Base for Video-Language Retrieval. CVPR, 2023. [[paper]](https://openaccess.thecvf.com/content/CVPR2023/papers/Pei_CLIPPING_Distilling_CLIP-Based_Models_With_a_Student_Base_for_Video-Language_CVPR_2023_paper.pdf) - `[Li et al. CVPR23]` SViTT: Temporal Learning of Sparse Video-Text Transformers. CVPR, 2023. [[paper]](https://openaccess.thecvf.com/content/CVPR2023/papers/Li_SViTT_Temporal_Learning_of_Sparse_Video-Text_Transformers_CVPR_2023_paper.pdf) [[code]](http://svcl.ucsd.edu/projects/svitt/) - `[Wu et al. CVPR23]` Cap4Video: What Can Auxiliary Captions Do for Text-Video Retrieval. CVPR, 2023. [[paper]](https://openaccess.thecvf.com/content/CVPR2023/papers/Wu_Cap4Video_What_Can_Auxiliary_Captions_Do_for_Text-Video_Retrieval_CVPR_2023_paper.pdf) [[code]](https://github.com/whwu95/Cap4Video) - `[Ko et al. CVPR23]` MELTR: Meta Loss Transformer for Learning to Fine-tune Video Foundation Models. CVPR, 2023. [[paper]](https://openaccess.thecvf.com/content/CVPR2023/papers/Ko_MELTR_Meta_Loss_Transformer_for_Learning_To_Fine-Tune_Video_Foundation_CVPR_2023_paper.pdf) [[code]](https://github.com/mlvlab/) - `[Wang et al. CVPR23]` All in One: Exploring Unified Video-Language Pre-Training. CVPR, 2023. [[paper]](https://openaccess.thecvf.com/content/CVPR2023/html/Wang_All_in_One_Exploring_Unified_Video-Language_Pre-Training_CVPR_2023_paper.html) [[code]](https://github.com/showlab/all-in-one) - `[Girdhar et al. CVPR23]` IMAGEBIND: One Embedding Space To Bind Them All. CVPR, 2023. [[paper]](https://openaccess.thecvf.com/content/CVPR2023/papers/Girdhar_ImageBind_One_Embedding_Space_To_Bind_Them_All_CVPR_2023_paper.pdf) [[code]](https://facebookresearch.github.io/ImageBind) - `[Huang et al. CVPR23]` VoP: Text-Video Co-Operative Prompt Tuning for Cross-Modal Retrieval. CVPR, 2023. [[paper]](https://openaccess.thecvf.com/content/CVPR2023/html/Huang_VoP_Text-Video_Co-Operative_Prompt_Tuning_for_Cross-Modal_Retrieval_CVPR_2023_paper.html) [[code]](https://github.com/bighuang624/VoP) - `[Li et al. CVPR23]` LAVENDER: Unifyi
Excerpt of 29,028 characters
Read on GitHub20
16
1
1
1
1
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:0b16e41435937fcc, desc:curated list