Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Variants of Vision Transformer and its downstream tasks
| Date | Stars |
|---|---|
| 2026-07-24 | 257 |
| 2026-07-25 | 257 |
| 2026-07-28 | 257 |
| 2026-07-30 | 257 |
| 2026-08-06 | 257 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# __Awesome Vision Transformer Collection__ __Variants of Vision Transformer and Vision Transformer for Downstream Tasks__ author: Runwei Guan affiliation: University of Liverpool / JITRI-Institute of Deep Perception Technology email: [email protected] / [email protected] / [email protected] ## Image Backbone * Vision Transformer [paper](https://arxiv.org/abs/2010.11929) [code](https://github.com/google-research/vision_transformer) * Swin Transformer [paper](https://arxiv.org/abs/2103.14030) [code](https://github.com/microsoft/Swin-Transformer) * Swin Transformer V2: Scaling Up Capacity and Resolution [paper](https://arxiv.org/abs/2111.09883) [code](https://github.com/microsoft/Swin-Transformer) * DVT [paper](https://arxiv.org/abs/2105.15075) [code](https://github.com/blackfeather-wang/Dynamic-Vision-Transformer) * PVT [paper](https://arxiv.org/abs/2102.12122) [code](https://github.com/whai362/PVT) * Lite Vision Transformer: LVT [paper](https://arxiv.org/abs/2112.10809) * PiT [paper](https://arxiv.org/abs/2103.16302) [code](https://github.com/naver-ai/pit) * Twins [paper](https://arxiv.org/abs/2104.13840) [code](https://github.com/Meituan-AutoML/Twins) * TNT [paper](https://arxiv.org/abs/2103.00112) [code](https://github.com/lucidrains/transformer-in-transformer) * Mobile-ViT [paper](https://arxiv.org/abs/2110.02178?context=cs.LG) [code](https://github.com/chinhsuanwu/mobilevit-pytorch) * Cross-ViT [paper](https://openaccess.thecvf.com/content/ICCV2021/html/Chen_CrossViT_Cross-Attention_Multi-Scale_Vision_Transformer_for_Image_Classification_ICCV_2021_paper.html) [code](https://github.com/IBM/CrossViT) * LeViT [paper](https://arxiv.org/pdf/2104.01136.pdf) [code](https://github.com/facebookresearch/LeViT) * ViT-Lite [paper](https://arxiv.org/pdf/2104.05704.pdf) * Refiner [paper](https://arxiv.org/pdf/2106.03714.pdf) [code](https://github.com/zhoudaquan/Refiner_ViT) * DeepViT [paper](https://arxiv.org/pdf/2103.11886.pdf) [code](https://github.com/zhoudaquan/dvit_repo) * CaiT [paper](https://arxiv.org/pdf/2103.17239.pdf) [code](https://github.com/facebookresearch/deit) * LV-ViT [paper](https://arxiv.org/pdf/2104.10858.pdf) [code](https://github.com/zihangJiang/TokenLabeling) * DeiT [paper](https://arxiv.org/pdf/2012.12877.pdf) [code](https://github.com/facebookresearch/deit) * CeiT [paper](https://arxiv.org/pdf/2103.11816.pdf) [code](https://github.com/rishikksh20/CeiT-pytorch) * BoTNet [paper](https://arxiv.org/abs/2101.11605) * ViTAE [paper](https://arxiv.org/abs/2106.03348) * Visformer: The Vision-Friendly Transformer [paper](https://openaccess.thecvf.com/content/ICCV2021/html/Chen_Visformer_The_Vision-Friendly_Transformer_ICCV_2021_paper.html) [code](https://github.com/danczs/Visformer) * Bootstrapping ViTs: Towards Liberating Vision Transformers from Pre-training [paper](https://arxiv.org/abs/2112.03552) * AdaViT: Adaptive Tokens for Efficient Vision Transformer [paper](https://arxiv.org/abs/2112.07658) * Improved Multiscale Vision Transformers for Classification and Detection [paper](https://arxiv.org/abs/2112.01526) * Multi-Scale Vision Longformer: A New Vision Transformer for High-Resolution Image Encoding [paper](https://openaccess.thecvf.com/content/ICCV2021/html/Zhang_Multi-Scale_Vision_Longformer_A_New_Vision_Transformer_for_High-Resolution_Image_ICCV_2021_paper.html) * Towards End-to-End Image Compression and Analysis with Transformers [paper](https://arxiv.org/abs/2112.09300) * MPViT: Multi-Path Vision Transformer for Dense Prediction [paper](https://arxiv.org/abs/2112.11010) * Lite Vision Transformer with Enhanced Self-Attention [paper](https://arxiv.org/abs/2112.10809) * PolyViT: Co-training Vision Transformers on Images, Videos and Audio [paper](https://arxiv.org/abs/2111.12993) * MIA-Former: Efficient and Robust Vision Transformers via Multi-grained Input-Adaptation [paper](https://arxiv.org/abs/2112.11542) * ELSA: Enhanced Local Self-Attention for Vision Transformer [paper](https://arxiv.or
Excerpt of 33,822 characters
Read on GitHubWould you bet a product on this? Bounded 0–100 and slow moving.
matched fp:148fd424f4ca6b9f, topic:computer-vision, topic:pose-estimation, readme:image classification
matched fp:148fd424f4ca6b9f, topic:deep-learning, readme:pre-training
matched fp:148fd424f4ca6b9f, topic:explainable-ai
matched fp:148fd424f4ca6b9f, topic:awesome