Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Collect some papers about transformer with vision. Awesome Transformer with Computer Vision (CV)
| Date | Stars |
|---|---|
| 2026-07-24 | 3589 |
| 2026-07-25 | 3589 |
| 2026-07-28 | 3589 |
| 2026-07-30 | 3589 |
| 2026-07-31 | 3590 |
| 2026-08-06 | 3590 |
Today
— stars today
This week
+1 stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.03%/day
# Awesome Visual-Transformer [](https://github.com/sindresorhus/awesome) Collect some Transformer with Computer-Vision (CV) papers. If you find some overlooked papers, please open issues or pull requests (recommended). ## Papers ### Transformer original paper - [Attention is All You Need](https://arxiv.org/abs/1706.03762) (NIPS 2017) ### Technical blog - [English Blog] Transformers in Vision [[Link](https://davide-coccomini.medium.com/)] - [Chinese Blog] 3W字长文带你轻松入门视觉transformer [[Link](https://zhuanlan.zhihu.com/p/308301901)] - [Chinese Blog] Vision Transformer 超详细解读 (原理分析+代码解读) [[Link](https://zhuanlan.zhihu.com/p/348593638)] ### Survey - Multimodal learning with transformers: A survey (IEEE TPAMI) [[paper](https://arxiv.org/abs/2206.06488)] - 2023.05.11 - A Survey of Visual Transformers [[paper](https://arxiv.org/abs/2111.06091)] - 2021.11.30 - Transformers in Vision: A Survey [[paper](https://arxiv.org/abs/2101.01169)] - 2021.02.22 - A Survey on Visual Transformer [[paper](https://arxiv.org/abs/2012.12556)] - 2021.1.30 - A Survey of Transformers [[paper](https://arxiv.org/abs/2106.04554)] - 2020.6.09 ### arXiv papers - **[Superpoint Transformer]** Efficient 3D Semantic Segmentation with Superpoint Transformer [[paper](https://arxiv.org/abs/2306.08045)] [[code](https://github.com/drprojects/superpoint_transformer)] - Understanding Gaussian Attention Bias of Vision Transformers Using Effective Receptive [[paper](https://arxiv.org/abs/2305.04722)] - **[FocusedDecoder]** Focused Decoding Enables 3D Anatomical Detection by Transformers [[paper](https://arxiv.org/abs/2207.10774v4)] [[code](https://github.com/bwittmann/transoar)] - **[TAG]** TAG: Boosting Text-VQA via Text-aware Visual Question-answer Generation [[paper](https://arxiv.org/abs/2208.01813)] [[code](https://github.com/HenryJunW/TAG)] - **[FastMETRO]** Cross-Attention of Disentangled Modalities for 3D Human Mesh Recovery with Transformers [[paper](https://arxiv.org/abs/2207.13820)] [[code](https://github.com/postech-ami/FastMETRO)] - BatchFormer: Learning to Explore Sample Relationships for Robust Representation Learning [[paper](https://arxiv.org/abs/2203.01522)] [[code](https://github.com/zhihou7/BatchFormer)] - **[RelViT]** RelViT: Concept-guided Vision Transformer for Visual Relational Reasoning [[paper]](https://arxiv.org/pdf/2204.11167.pdf) [[code]](https://github.com/NVlabs/RelViT) - **[MViTv2]** Improved Multiscale Vision Transformers for Classification and Detection [[paper](https://arxiv.org/pdf/2112.01526.pdf)] [[code](https://github.com/facebookresearch/mvit)] - DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection [[paper](https://arxiv.org/pdf/2203.03605.pdf)] [[code](https://github.com/IDEACVR/DINO)] - Three things everyone should know about Vision Transformers [[paper](https://arxiv.org/pdf/2203.09795.pdf)] - **[DeiT III]** DeiT III: Revenge of the ViT [[paper](https://arxiv.org/pdf/2204.07118.pdf)] - **[DaViT]** DaViT: Dual Attention Vision Transformers [[paper](https://arxiv.org/pdf/2204.03645.pdf)] [[code](https://github.com/dingmyu/davit)] - **[CoFormer]** Collaborative Transformers for Grounded Situation Recognition [[paper](https://arxiv.org/abs/2203.16518)] [[code](https://github.com/jhcho99/CoFormer)] - **[GSRTR]** Grounded Situation Recognition with Transformers [[paper](https://arxiv.org/abs/2111.10135)] [[code](https://github.com/jhcho99/gsrtr)] - **[MaxViT]** MaxViT: Multi-Axis Vision Transformer [[paper]](https://arxiv.org/abs/2204.01697) - **[V2X-ViT]** V2X-ViT: Vehicle-to-Everything Cooperative Perception with Vision Transformer [[paper]](https://arxiv.org/abs/2203.10638) - **[MemMC-MAE]** Unsupervised Anomaly Detection in Medical Images with a Memory-augmented Multi-level Cross-attentional Masked Autoencoder [[paper](https://arxiv.org/abs/2203.11725)] [[code](https://github.c
Excerpt of 73,483 characters
Read on GitHubDingkang Liang · Huazhong University of Science & Technology · China
215
2
2
Tianrui Guan
2
2
2
2
China
2
Bin Yan · ByteDance · United States
2
2
Damien ROBERT · University of Zurich · Switzerland
1
1
1
1
Bolin Ni · Institute of Automation, Chinese Academy of Sciences · China
1
1
1
Hao Luo · Zhejiang University, China · China
1
Fawaz Sammani · Vrije Universiteit Brussel · Belgium
1
Zhengzhong Tu · @TAMU @Google @Google-Research @UTAustin
1
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:d5951dbc10d1650c, topic:transformer
matched fp:d5951dbc10d1650c, desc:computer vision, readme:computer vision, readme:object detection