Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Explainability for Vision Transformers
| Date | Stars |
|---|---|
| 2026-07-24 | 1097 |
| 2026-07-25 | 1097 |
| 2026-07-28 | 1097 |
| 2026-07-30 | 1097 |
| 2026-08-06 | 1097 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# Explainability for Vision Transformers (in PyTorch)
This repository implements methods for explainability in Vision Transformers.
See also https://jacobgil.github.io/deeplearning/vision-transformer-explainability
## Currently implemented:
- Attention Rollout.
- Gradient Attention Rollout for class specific explainability.
*This is our attempt to further build upon and improve Attention Rollout.*
- TBD Attention flow is work in progress.
Includes some tweaks and tricks to get it working:
- Different Attention Head fusion methods,
- Removing the lowest attentions.
## Usage
- From code
``` python
from vit_grad_rollout import VITAttentionGradRollout
model = torch.hub.load('facebookresearch/deit:main',
'deit_tiny_patch16_224', pretrained=True)
grad_rollout = VITAttentionGradRollout(model, discard_ratio=0.9, head_fusion='max')
mask = grad_rollout(input_tensor, category_index=243)
```
- From the command line:
```
python vit_explain.py --image_path <image path> --head_fusion <mean, min or max> --discard_ratio <number between 0 and 1> --category_index <category_index>
```
If category_index isn't specified, Attention Rollout will be used,
otherwise Gradient Attention Rollout will be used.
Notice that by default, this uses the 'Tiny' model from [Training data-efficient image transformers & distillation through attention](https://arxiv.org/abs/2012.12877)
hosted on torch hub.
## Where did the Transformer pay attention to in this image?
| Image | Vanilla Attention Rollout | With discard_ratio+max fusion |
| -------------------------|-------------------------|------------------------- |
|  |  | 
 |  |  |
 |  |  |
 |  |  |
## Gradient Attention Rollout for class specific explainability
The Attention that flows in the transformer passes along information belonging to different classes.
Gradient roll out lets us see what locations the network paid attention too,
but it tells us nothing about if it ended up using those locations for the final classification.
We can multiply the attention with the gradient of the target class output, and take the average among the attention heads (while masking out negative attentions) to keep only attention that contributes to the target category (or categories).
### Where does the Transformer see a Dog (category 243), and a Cat (category 282)?
 
### Where does the Transformer see a Musket dog (category 161) and a Parrot (category 87):
 
## Tricks and Tweaks to get this working
### Filtering the lowest attentions in every layer
`--discard_ratio <value between 0 and 1>`
Removes noise by keeping the strongest attentions.
Results for dIfferent values:
 
### Different Attention Head Fusions
The Attention Rollout method suggests taking the average attention accross the attention heads,
but emperically it looks like taking the Minimum value, Or the Maximum value combined with --discard_ratio, works better.
` --head_fusion <mean, min or max>`
| Image | Mean Fusion | Min Fusion |
| -------------------------|-------------------------|------------------------- |
 |  | 
## References
-Excerpt of 4,571 characters
Read on GitHubJacob Gildenblat · Israel
19
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:ac196137bf7ab606, topic:deep-learning, topic:pytorch
matched fp:ac196137bf7ab606, topic:transformer
matched fp:ac196137bf7ab606, topic:explainable-ai