Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Course Materials for Interpretability of Large Language Models (0368.4264) at Tel Aviv University
| Date | Stars |
|---|---|
| 2026-07-31 | 333 |
| 2026-08-06 | 334 |
Today
+1 stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# Interpretability of Large Language Models (0368.4264) This repository contains materials for the **Interpretability of Large Language Models** course (0368.4264) at Tel Aviv University. It is a graduate-level, active-learning course in which students learn about interpretability of LLMs in the style of a collaborative research group. The course is structured around weekly paper readings, in-class discussions, role-playing, and hands-on exercises.[^1] Students are assumed to have prior background in natural language processing and machine learning. In this repository, you will find: * Schedule and reading lists * Coding exercises and challenges The course was developed by Dr. Mor Geva and Daniela Gottesman at Tel Aviv University. We also thank Amit Elhelo, Or Shafran, and Yoav Gur-Arieh for their contributions. We share these materials and hope they serve as a useful resource for anyone curious about or working on the interpretability of large language models. [^1]: The course format draws inspiration from the [paper-reading seminar by Alec Jacobson and Colin Raffel](https://colinraffel.com/blog/role-playing-seminar.html) and [The Science of Large Language Models course by Robin Jia](https://robinjia.github.io/classes/fall2024-csci699.html). ## Schedule and materials The schedule is subject to minor changes. Week | Date | Topic and papers | Practicum | ------- | ------ | ------- | ------- | | 1 | Oct 26 | **Introduction and role assignments**<br>Background and NLP refresher | [Exercise](https://colab.research.google.com/drive/1hAsTsbnwb6jZkfRyRPokTOWnoK34UdB1?usp=sharing) [Solution](https://colab.research.google.com/drive/1kTcuXuRjkzpmuacW-MSNK4bMKJVef5Sy?usp=sharing)| | 2 | Nov 2 | **Probing**<br>Main paper 1: [Language Models Represent Space and Time](https://arxiv.org/abs/2310.02207)<br>Main paper 2: [A Structural Probe for Finding Syntax in Word Representations](https://aclanthology.org/N19-1419/)<br>Bonus papers:<br> * [Not All Language Model Features Are One-Dimensionally Linear](https://arxiv.org/abs/2405.14860) |[Exercise](https://colab.research.google.com/drive/1NSWt5QutY-4hJtY2jG23LkUfY7mEFuub?usp=sharing) [Solution](https://colab.research.google.com/drive/13SEeFqWmjXZjSAWxk4YYd0P282G7-e2R?usp=sharing) | | 3 | Nov 9 | **Inspecting representations**<br>Main paper 1: [Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models](https://arxiv.org/abs/2401.06102)<br>Main paper 2: [Language Model Inversion](https://arxiv.org/abs/2311.13647)<br>Bonus papers:<br> * [SelfIE: Self-Interpretation of Large Language Model Embeddings](https://arxiv.org/abs/2403.10949)<br> * [LatentQA: Teaching LLMs to Decode Activations Into Natural Language](https://arxiv.org/abs/2412.08686)<br> * [logit lens](https://www.lesswrong.com/posts/AcKRB8wDpdaN6v6ru/interpreting-gpt-the-logit-lens) |[Exercise](https://colab.research.google.com/drive/15W76ULQvgayXcXzMJ9MBeijrkDAq1k80?usp=sharing) [Solution](https://colab.research.google.com/drive/1Be42DpybrBxO3I-vZVyVuwo_acgIJucQ?usp=sharing) | | 4 | Nov 16 | **Attention heads**<br>Main paper 1: [Inferring Functionality of Attention Heads from their Parameters](https://arxiv.org/abs/2412.11965)<br>Main paper 2: [Talking Heads: Understanding Inter-layer Communication in Transformer Language Models](https://arxiv.org/abs/2406.09519)<br>Bonus papers:<br> * [In-context Learning and Induction Heads](https://transformer-circuits.pub/2022/in-context-learning-and-induction-heads/index.html)<br> * [Attention Heads of Large Language Models: A Survey](https://arxiv.org/abs/2409.03752)<br> * [Analyzing Transformers in Embedding Space](https://arxiv.org/abs/2209.02535) |[Exercise](https://colab.research.google.com/drive/1KKTW04LBoNNnO-hF24GOOCgUs7QBiahW?usp=sharing) [Solution](https://colab.research.google.com/drive/1VpV2WM7TdCJPnhvBttaw0laBnjvpSfN3?usp=sharing)| | 5 | Nov 23 | **MLP layers**<br>Main paper 1: [Transformer Feed-Forward Layers Are Key-Value Memories](https://arxi
Excerpt of 10,008 characters
Read on GitHubWould you bet a product on this? Bounded 0–100 and slow moving.
matched fp:09aef3008768557d, desc:interpretability