Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
A curated list of LLM Interpretability related material - Tutorial, Library, Survey, Paper, Blog, etc..
| Date | Stars |
|---|---|
| 2026-07-31 | 307 |
| 2026-08-06 | 307 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# Awesome-LLM-Interpretability
A curated list of LLM Interpretability related material.
## ToC
- [Tutorial](#tutorial)
- [History](#history)
- [Code](#code)
- [Library](#library)
- [Codebase](#codebase)
- [Survey](#survey)
- [Video](#video)
- [Paper & Blog](#paper--blog)
- [By Source](#by-source)
- [By Topic](#by-topic)
- [Tools/Techniques/Methods](#toolstechniquesmethods)
- [General](#general)
- [Embedding Projection](#embedding-projection)
- [Probing](#probing)
- [Causal Intervention](#causal-intervention)
- [Automation](#automation)
- [Sparse Coding](#sparse-coding)
- [Visualization](#visualization)
- [Translation](#translation)
- [Evaluation/Dataset/Benchmark](#evaluationdatasetbenchmark)
- [Task Solving/Function/Ability](#task-solvingfunctionability)
- [General](#general-1)
- [Reasoning](#reasoning)
- [Function](#function)
- [Arithmetic Ability](#arithmetic-ability)
- [In-context Learning](#in-context-learning)
- [Factual Knowledge](#factual-knowledge)
- [Multilingual/Crosslingual](#multilingualcrosslingual)
- [Multimodal](#multimodal)
- [Component](#component)
- [General](#general-2)
- [Attention](#attention)
- [MLP/FFN](#mlpffn)
- [Neuron](#neuron)
- [Learning Dynamics](#learning-dynamics)
- [General](#general-3)
- [Phase Transition/Grokking](#phase-transitiongrokking)
- [Fine-tuning](#fine-tuning)
- [Feature Representation/Probing-based](#feature-representationprobing-based)
- [General](#general-4)
- [Linearity](#linearity)
- [Application](#application)
- [Inference-Time Intervention/Activation Steering](#inference-time-interventionactivation-steering)
- [Knowledge/Model Editing](#knowledgemodel-editing)
- [Hallucination](#hallucination)
- [Pruning/Redundancy Analysis](#pruningredundancy-analysis)
## Tutorial
* **Concrete Steps to Get Started in Transformer Mechanistic Interpretability** [[Neel Nanda's blog]](https://www.neelnanda.io/mechanistic-interpretability/getting-started)
* **Mechanistic Interpretability Quickstart Guide** [[Neel Nanda's blog]](https://www.neelnanda.io/mechanistic-interpretability/getting-started)
* **ARENA Mechanistic Interpretability Tutorials by Callum McDougall** [[website]](https://arena-ch1-transformers.streamlit.app/)
* **200 Concrete Open Problems in Mechanistic Interpretability: Introduction by Neel Nanda** [[AlignmentForum]](https://www.alignmentforum.org/s/yivyHaCAmMJ3CqSyj)
* **Transformer-specific Interpretability** [[EACL 2023 Tutorial]](https://projects.illc.uva.nl/indeep/tutorial/)
## History
* **Mechanistic?** [[BlackBoxNLP workshop at EMNLP 2024]](https://arxiv.org/abs/2410.09087)
* This paper explores the multiple definitions and uses of "mechanistic interpretability," tracing its evolution in NLP research and revealing a critical divide within the interpretability community.
## Platform
* **Neuronpedia** [[website]](https://www.neuronpedia.org/)
* Accelerating researchers for Sparse Autoencoders (SAEs) by hosting models, feature dashboards, data visualizations, tooling, and more
## Code
### Library
* **TransformerLens** [[github]](https://github.com/neelnanda-io/TransformerLens)
* A library for mechanistic interpretability of GPT-style language models
* **SAELens** [[github]](https://github.com/jbloomAus/SAELens)
* Training and analyzing sparse autoencoders on Language Models
* **NNSight** [[github]](https://github.com/ndif-team/nnsight)
* The nnsight package enables interpreting and manipulating the internals of deep learned models.
* **CircuitsVis** [[github]](https://github.com/alan-cooney/CircuitsVis)
* Mechanistic Interpretability visualizations
* **baukit** [[github]](https://github.com/davidbau/baukit)
* Contains some methods for tracing and editing internal activations in a network.
* **transformer-debugger** [[github]](https://githubExcerpt of 48,026 characters
Read on GitHubWould you bet a product on this? Bounded 0–100 and slow moving.
matched fp:4011941f58abd018, llm:Repository description: 'A curated list of LLM Interpretability related material - Tutorial, Library, Survey, Paper, Blog, etc.' (Awesome list of LLM interpretability resources).
matched fp:4011941f58abd018, llm:Repository description: 'A curated list of LLM Interpretability related material - Tutorial, Library, Survey, Paper, Blog, etc.' (Awesome list of LLM interpretability resources).
matched fp:4011941f58abd018, llm:Repository description: 'A curated list of LLM Interpretability related material - Tutorial, Library, Survey, Paper, Blog, etc.' (Awesome list of LLM interpretability resources).