Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Awesome-LLM-KV-Cache: A curated list of 📙Awesome LLM KV Cache Papers with Codes.
| Date | Stars |
|---|---|
| 2026-07-31 | 461 |
| 2026-08-06 | 462 |
Today
+1 stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
<!--  --> <!-- <div align='center'> <img src=https://cdn.rawgit.com/sindresorhus/awesome/d7305f38d29fed78fa85652e3a63e154dd8e8829/media/badge.svg > <img src=https://img.shields.io/github/downloads/Zefan-Cai/Awesome-LLM-KV-Cache/total?color=ccf&label=downloads&logo=github&logoColor=lightgrey > <img src=https://img.shields.io/github/forks/Zefan-Cai/Awesome-LLM-KV-Cache.svg?style=social > <img src=https://img.shields.io/github/stars/Zefan-Cai/Awesome-LLM-KV-Cache.svg?style=social > <img src=https://img.shields.io/github/watchers/Zefan-Cai/Awesome-LLM-KV-Cache.svg?style=social > <img src=https://img.shields.io/badge/Release-v1.6-brightgreen.svg > <img src=https://img.shields.io/badge/License-GPLv3.0-turquoise.svg > </div> --> ## 📒Introduction Awesome-LLM-KV-Cache: A curated list of [📙Awesome LLM KV Cache Papers with Codes](#paperlist). This repository is for personal use of learning and classifying the burning KV Cache related papers! ## ©️Citations ## 📖Contents <div id="paperlist"></div> * 📖[Trending Inference Topics](#Trending-Inference-Topics)🔥🔥🔥 * 📖[KV Cache Compression](#KV-Cache-Compression)🔥🔥 * 📖[KV Cache Merge](#KV-Cache-Merge)🔥🔥 * 📖[Budget Allocation](#Budget-Allocation)🔥 * 📖[Query-Aware KV Retrieval](#Query-Aware-KV-Retrieval)🔥 * 📖[Cross-Layer KV Cache Utilization](#Cross-Layer-KV-Cache-Utilization)🔥 * 📖[KV Cache Quantization](#KV-Cache-Quantization)🔥 * 📖[Low-Rank KV Cache Decomposition](#Low-Rank-KV-Cache-Decomposition)🔥 * 📖[Observation](#Observation)🔥🔥 * 📖[Evaluation](#Evaluation)🔥 * 📖[Systems](#Systems) * 📖[Others](#Others) ### 📖Trending Inference Topics ([©️back👆🏻](#paperlist)) <div id="Trending-Inference-Topics"></div> |Date|Title|Paper|Code|Recom|Comment| |:---:|:---:|:---:|:---:|:---:|:---:| |2024.05| 🔥🔥🔥[DeepSeek-V2] DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model(@DeepSeek-AI)|[[pdf]](https://arxiv.org/pdf/2405.04434) | [[DeepSeek-V2]](https://github.com/deepseek-ai/DeepSeek-V2) | ⭐️⭐️⭐️ | |2024.05|🔥🔥🔥[YOCO] You Only Cache Once: Decoder-Decoder Architectures for Language Models(@Microsoft)| [[pdf]](https://arxiv.org/pdf/2405.05254) | [[unilm-YOCO]](https://github.com/microsoft/unilm/tree/master/YOCO)  |⭐️⭐️⭐️ | |2024.06|🔥🔥[**Mooncake**] Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving(@Moonshot AI) |[[pdf]](https://github.com/kvcache-ai/Mooncake/blob/main/Mooncake-v1.pdf) | [[Mooncake]](https://github.com/kvcache-ai/Mooncake) |⭐️⭐️⭐️ | |2024.07|🔥🔥🔥[**FlashAttention-3**] FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision(@TriDao etc) |[[pdf]](https://tridao.me/publications/flash3/flash3.pdf)|[[flash-attention]](https://github.com/Dao-AILab/flash-attention) |⭐️⭐️⭐️ | |2024.07|🔥🔥🔥[**MInference 1.0**] MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention(@Microsoft) |[[pdf]](https://arxiv.org/pdf/2407.02490)|[[MInference 1.0]](https://github.com/microsoft/MInference) |⭐️⭐️⭐️ | ### LLM KV Cache Compression ([©️back👆🏻](#paperlist)) <div id="#KV-Cache-Compression"></div> |Date|Title|Paper|Code|Recom|Comment| |:---:|:---:|:---:|:---:|:---:|:---:| |2023.06| 🔥🔥[**H2O**] H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models|[[pdf]](https://arxiv.org/abs/2306.14048) | [[H2O]](https://github.com/FMInference/H2O) ![](https://img.shields.io/github/stars/FMInference/H2O.svg?style=s
Excerpt of 19,179 characters
Read on GitHubWould you bet a product on this? Bounded 0–100 and slow moving.
matched fp:cd8008cb5550f35c, topic:llm
matched fp:cd8008cb5550f35c, name:kv cache, desc:kv cache