Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
This repository collects all relevant resources about interpretability in LLMs
| Date | Stars |
|---|---|
| 2026-07-31 | 402 |
| 2026-08-05 | 402 |
| 2026-08-06 | 402 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# Awesome Interpretability in Large Language Models The area of interpretability in large language models (LLMs) has been growing rapidly in recent years. This repository tries to collect all relevant resources to help beginners quickly get started in this area and help researchers to keep up with the latest research progress. This is an active repository and welcome to open a new issue if I miss any relevant resources. If you have any questions or suggestions, please feel free to contact me via email: `[email protected]`. --- <font size=5><center><b> Table of Contents </b> </center></font> - [Awesome Interpretability Libraries](#awesome-interpretability-libraries) - [Awesome Interpretability Blogs & Videos](#awesome-interpretability-blogs--videos) - [Awesome Interpretability Tutorials](#awesome-interpretability-tutorials) - [Awesome Interpretability Forums](#awesome-interpretability-forums) - [Awesome Interpretability Tools](#awesome-interpretability-tools) - [Awesome Interpretability Programs](#awesome-interpretability-programs) - [Awesome Interpretability Papers](#awesome-interpretability-papers) - [Survey Papers](#survey-papers) - [Position Papers](#position-papers) - [Interpretable Analysis of LLMs](#interpretable-analysis-of-llms) - [SAE, Dictionary Learning and Superposition](#sae-dictionary-learning-and-superposition) - [Interpretability in Vision LLMs](#interpretability-in-vision-llms) - [Benchmarking Interpretability](#benchmarking-interpretability) - [Enhancing Interpretability](#enhancing-interpretability) - [Others](#others) - [Other Awesome Interpretability Resources](#other-awesome-interpretability-resources) --- # Awesome Interpretability Libraries -  [**TransformerLens**](https://github.com/TransformerLensOrg/TransformerLens): A Library for Mechanistic Interpretability of Generative Language Models. ([Doc](https://transformerlensorg.github.io/TransformerLens/), [Tutorial](https://arena3-chapter1-transformer-interp.streamlit.app/[1.2]_Intro_to_Mech_Interp), [Demo](https://colab.research.google.com/github/neelnanda-io/TransformerLens/blob/main/demos/Main_Demo.ipynb)) -  [**nnsight**](https://github.com/ndif-team/nnsight): enables interpreting and manipulating the internals of deep learned models. ([Doc](https://nnsight.net/documentation/), [Tutorial](https://nnsight.net/tutorials/), [Paper](https://arxiv.org/abs/2407.14561)) -  [**SAE Lens**](https://github.com/jbloomAus/SAELens): train and analyse SAE. ([Doc](https://jbloomaus.github.io/SAELens/), [Tutorial](https://github.com/jbloomAus/SAELens/tree/main/tutorials), [Blog](https://www.lesswrong.com/posts/f9EgfLSurAiqRJySD/open-source-sparse-autoencoders-for-all-residual-stream)) -  [**EleutherAI: sae**](https://github.com/EleutherAI/sae): train SAE on very large model based on the method and released code of the [openAI SAE paper](https://arxiv.org/abs/2406.04093v1) -  [**Automatic Circuit DisCovery**](https://github.com/ArthurConmy/Automatic-Circuit-Discovery): automatically build circuit for mechanistic interpretability. ([Paper](https://arxiv.org/pdf/2304.14997), [Demo](https://colab.research.google.com/github/ArthurConmy/Automatic-Circuit-Discovery/blob/main/notebooks/colabs/ACDC_Main_Demo.ipynb)) -  [**Pyvene**](https://github.com/stanfordnlp/pyvene): A Library for Understanding and Improving PyTorch Models via Interventions. ([Paper](https://arxiv.org/pdf/2403.07809), [Demo](https://colab.research.google.com/github/stanfordnlp/pyvene/blob/main/pyvene
Excerpt of 61,721 characters
Read on GitHubWould you bet a product on this? Bounded 0–100 and slow moving.
matched fp:c567de284aefe72f, name:interpretability, desc:interpretability