Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Awesome Transformers (self-attention) in Computer Vision
| Date | Stars |
|---|---|
| 2026-07-24 | 271 |
| 2026-07-25 | 271 |
| 2026-07-28 | 271 |
| 2026-07-30 | 271 |
| 2026-07-31 | 271 |
| 2026-08-06 | 271 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# Awesome Visual Representation Learning with Transformers [](https://awesome.re) Awesome Transformers (self-attention) in Computer Vision ## About transformers - Attention Is All You Need, NeurIPS 2017 - Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, Illia Polosukhin - [[paper]](https://arxiv.org/abs/1706.03762) [[official code]](https://github.com/tensorflow/tensor2tensor) [[pytorch implementation]](https://github.com/jadore801120/attention-is-all-you-need-pytorch) - BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, NAACL 2019 - Jacob Devlin, Ming-Wei Chang, Kenton Lee, Kristina Toutanova - [[paper]](https://arxiv.org/abs/1810.04805) [[offficial code]](https://github.com/google-research/bert) [[huggingface/transformers]](https://github.com/huggingface/transformers) - Efficient Transformers: A Survey, arXiv 2020 - Yi Tay, Mostafa Dehghani, Dara Bahri, Donald Metzler - [[paper]](https://arxiv.org/abs/2009.06732) - A Survey on Visual Transformer, arXiv 2020 - Kai Han, Yunhe Wang, Hanting Chen, Xinghao Chen, Jianyuan Guo, Zhenhua Liu, Yehui Tang, An Xiao, Chunjing Xu, Yixing Xu, Zhaohui Yang, Yiman Zhang, Dacheng Tao - [[paper]](https://arxiv.org/abs/2012.12556) - Transformers in Vision: A Survey, arXiv 2021 - Salman Khan, Muzammal Naseer, Munawar Hayat, Syed Waqas Zamir, Fahad Shahbaz Khan, Mubarak Shah - [[paper]](https://arxiv.org/abs/2101.01169) ## Combining CNN with self-attention - Attention augmented convolutional networks, ICCV 2019, image classification - Irwan Bello, Barret Zoph, Ashish Vaswani, Jonathon Shlens, Quoc V. Le - [[paper]](https://arxiv.org/abs/1904.09925) [[pytorch implementation]](https://github.com/leaderj1001/Attention-Augmented-Conv2d) - Self-Attention Generative Adversarial Networks, ICML 2019, generative model(GANs) - Han Zhang, Ian Goodfellow, Dimitris Metaxas, Augustus Odena - [[paper]](https://arxiv.org/abs/1805.08318) [[official code]](https://github.com/heykeetae/Self-Attention-GAN) - Videobert: A joint model for video and language representation learning, ICCV 2019, video processing - Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, Cordelia Schmid - [[paper]](https://arxiv.org/abs/1904.01766) - Visual Transformers: Token-based Image Representation and Processing for Computer Vision, arXiv 2020, image classification - Bichen Wu, Chenfeng Xu, Xiaoliang Dai, Alvin Wan, Peizhao Zhang, Masayoshi Tomizuka, Kurt Keutzer, Peter Vajda - [[paper]](https://arxiv.org/abs/2006.03677) - Feature Pyramid Transformer, ECCV 2020, detection and segmentation - Dong Zhang, Hanwang Zhang, Jinhui Tang, Meng Wang, Xiansheng Hua, Qianru Sun - [[paper]](http://arxiv.org/abs/2007.09451) [[official code]](https://github.com/ZHANGDONG-NJUST/FPT) - Revisiting Stereo Depth Estimation From a Sequence-to-Sequence Perspective with Transformers, arXiv 2020, depth estimation - Zhaoshuo Li, Xingtong Liu, Francis X. Creighton, Russell H. Taylor, and Mathias Unberath - [[paper]](http://arxiv.org/abs/2011.02910) [[official code]](https://github.com/mli0603/stereo-transformer) - End-to-end Lane Shape Prediction with Transformers, arXiv 2020, lane detection - Ruijin Liu, Zejian Yuan, Tie Liu, Zhiliang Xiong - [[paper]](http://arxiv.org/abs/2011.04233) [[official code]](https://github.com/liuruijin17/LSTR) - Taming Transformers for High-Resolution Image Synthesis, arXiv 2020, image synthesis - Patrick Esser, Robin Rombach, Bjorn Ommer - [[paper]](http://arxiv.org/abs/2012.09841)[[official code]](https://github.com/CompVis/taming-transformers) - TransPose: Towards Explainable Human Pose Estimation by Transformer, arXiv 2020, pose estimation - Sen Yang, Zhibin Quan, Mu Nie, Wankou Yang - [[paper]](https://arxiv.org/abs/2012.14214) - End-to-End Video Instance Segmentation with Transformers, arXiv 2020, video instance segmentation - Yuqing Wang, Zha
Excerpt of 15,185 characters
Read on GitHubWould you bet a product on this? Bounded 0–100 and slow moving.
matched fp:b7c0d734078eb782, topic:computer-vision, desc:computer vision, readme:computer vision
matched fp:b7c0d734078eb782, topic:representation-learning, name:representation learning, readme:representation learning
matched fp:b7c0d734078eb782, topic:awesome-list