Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
awesome papers in LLM interpretability
| Date | Stars |
|---|---|
| 2026-07-31 | 625 |
| 2026-08-03 | 625 |
| 2026-08-06 | 625 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# Awesome Papers for Understanding LLM Mechanism This list focuses on understanding the internal mechanism of large language models (LLM). Works in this list are accepted by top conferences (e.g. ICML, NeurIPS, ICLR, ACL, EMNLP, NAACL), or written by top research institutions. Other paper lists focuses on [SAE](https://github.com/zepingyu0512/awesome-SAE) and [neuron](https://github.com/zepingyu0512/awesome-LLM-neuron). Paper recommendation (accepted by conferences): please contact [me](https://zepingyu0512.github.io/). ## Papers ### 2025 - [Locate-then-Merge: Neuron-Level Parameter Fusion for Mitigating Catastrophic Forgetting in Multimodal LLMs](https://arxiv.org/pdf/2505.16703) - \[EMNLP 2025\] \[2025.8\] \[multimodal\] \[model merging\] - [Back Attention: Understanding and Enhancing Multi-Hop Reasoning in Large Language Models](https://arxiv.org/pdf/2502.10835) - \[EMNLP 2025\] \[2025.8\] \[reasoning\] - [Neural Incompatibility: The Unbridgeable Gap of Cross-Scale Parametric Knowledge Transfer in Large Language Models](https://arxiv.org/pdf/2505.14436) - \[ACL 2025\] \[2025.5\] \[neuron\] - [Model Unlearning via Sparse Autoencoder Subspace Guided Projections](https://arxiv.org/abs/2505.24428) - \[ICML 2025 workshop\] \[2025.5\] \[SAE\] - [Lost in Multilinguality: Dissecting Cross-lingual Factual Inconsistency in Transformer Language Models](https://arxiv.org/pdf/2504.04264) - \[ACL 2025\] \[2025.4\] \[multilinguality\] - [On the Biology of a Large Language Model](https://transformer-circuits.pub/2025/attribution-graphs/biology.html) - \[Anthropic\] \[2025.3\] - [Taming Knowledge Conflicts in Language Models](https://www.arxiv.org/pdf/2503.10996) - \[ICML 2025\] \[2025.3\] \[knowledge\] \[hallucination\] \[superposition\] - [Circuit Tracing: Revealing Computational Graphs in Language Models](https://transformer-circuits.pub/2025/attribution-graphs/methods.html) - \[Anthropic\] \[2025.3\] - [The Mirage of Model Editing: Revisiting Evaluation in the Wild](https://arxiv.org/pdf/2502.11177) - \[ACL 2025\] \[2025.2\] \[model editing\] - [Towards Understanding Fine-Tuning Mechanisms of LLMs via Circuit Analysis](https://arxiv.org/abs/2502.11812) - \[ICML 2025\] \[2025.2\] \[circuit\] - [AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders](https://arxiv.org/pdf/2501.17148) - \[ICML 2025\] \[2025.1\] \[SAE\] ### 2024 - [Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models](https://arxiv.org/pdf/2412.17034) - \[ACL 2025\] \[2024.12\] \[safety\] - [Disentangling Memory and Reasoning Ability in Large Language Models](https://arxiv.org/pdf/2411.13504) - \[ACL 2025\] \[2024.11\] \[reasoning\] - [Can Knowledge Editing Really Correct Hallucinations?](https://arxiv.org/pdf/2410.16251) - \[ICLR 2025\] \[2024.10\] \[knowledge\] \[model editing\] - [Arithmetic without algorithms: Language models solve math with a bag of heuristics](https://arxiv.org/pdf/2410.21272) - \[ICLR 2025\] \[2024.10\] \[arithmetic\] - [Interpreting Arithmetic Mechanism in Large Language Models through Comparative Neuron Analysis](https://zepingyu0512.github.io/arithmetic-mechanism.github.io/) - \[EMNLP 2024\] \[2024.9\] \[neuron\] \[arithmetic\] \[fine-tune\] - [NNsight and NDIF: Democratizing Access to Open-Weight Foundation Model Internals](https://arxiv.org/pdf/2407.14561) - \[ICLR 2025\] \[2024.7\] - [Scaling and evaluating sparse autoencoders](https://arxiv.org/pdf/2406.04093) - \[OpenAI\] \[2024.6\] \[SAE\] - [BMIKE-53: Investigating Cross-Lingual Knowledge Editing with In-Context Learning](https://arxiv.org/pdf/2406.17764?) - \[ACL 2025\] \[2024.6\] \[model editing\] - [How do Large Language Models Learn In-Context? Query and Key Matrices of In-Context Heads are Two Towers for Metric Learning](https://zepingyu0512.github.io/in-context-mechanism.github.io/) - \[EMNLP 2024\] \[2024.6\] \[in-context learn
Excerpt of 20,354 characters
Read on GitHub121
Zifan Zheng
1
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:0ad3ac4c8c70da3f, desc:interpretability