Top AI Repos β open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
π Efficient implementations for emerging model architectures
| Date | Stars |
|---|---|
| 2026-07-24 | 5411 |
| 2026-07-25 | 5414 |
| 2026-07-28 | 5414 |
| 2026-07-30 | 5414 |
| 2026-08-06 | 5414 |
Today
β stars today
This week
β stars this week
This month
β stars this month
Momentum
15.0
growth rate 0.00%/day
<div align="center"> <img width="50%" alt="Flash Linear Attention" src="images/logo.png"> <br> [](https://huggingface.co/fla-hub) [](https://discord.gg/vDaJTmKNcS) </div> <p> π₯ Flash Linear Attention brings together hardware-efficient building blocks, training-ready layers, and components for modern sequence models, spanning linear attention, sparse attention, state space models, and hybrid LLM architectures. All implementations are platform-agnostic and verified on NVIDIA, AMD, and Intel hardware. Pull requests are welcome! </p> -------- * [News](#news) * [Models](#models) * [Installation](#installation) * [Usage](#usage) * [Token Mixing](#token-mixing) * [Fused Modules](#fused-modules) * [Generation](#generation) * [Hybrid Models](#hybrid-models) * [Training](#training) * [Evaluation](#evaluation) * [Benchmarks](#benchmarks) * [Citation](#citation) * [Star History](#star-history) * [Acknowledgements](#acknowledgements) ## News - [2026-07] π§± Add a [Gluon](https://triton-lang.org/main/getting-started/tutorials/gluon/) backend for [AttnRes](fla/ops/attnres). - [2026-07] π Add [FlashQLA](https://github.com/QwenLM/FlashQLA) backend for [Gated DeltaNet](fla/ops/gated_delta_rule). - [2026-06] π Add Parallax implementation to `fla` ([paper](https://arxiv.org/abs/2605.29157)). - [2026-06] π§± Add Wall attention implementation to `fla` ([blog](https://blog.tilderesearch.com/blog/wall-attn)). - [2026-05] πͺ Add Gated DeltaNet 2 (GDN-2) implementation to `fla` ([paper](https://arxiv.org/abs/2605.22791)). - [2026-05] π¦ Add Raven implementation to `fla` ([repo](https://github.com/goombalab/raven)). - [2026-05] π Add [YOCO](https://arxiv.org/abs/2405.05254) (You Only Cache Once) implementation to `fla`. - [2026-05] β‘ Add fused [AttnRes](fla/ops/attnres) support to `fla` ([paper](https://arxiv.org/abs/2603.15031)). - [2026-04] π Add Mamba3 implementation to `fla` ([paper](https://arxiv.org/abs/2603.15569)). - [2026-04] π§± Add [MoBA](https://arxiv.org/abs/2502.13189) (Mixture of Block Attention) implementation to `fla`, with [FlashMoBA](https://github.com/mit-han-lab/flash-moba) backend support. - [2026-04] π§± Add [TileLang](https://github.com/tile-ai/tilelang) backend support for selected kernels. - [2026-04] π― Add [GPT-OSS](https://openai.com/index/introducing-gpt-oss/)-style attention sink support to `fla`'s attention kernels. - [2026-03] π Add [Context Parallel](fla/ops/cp/README.md) support for KDA and GDN, enabling efficient distributed training across sequence dimension. - [2025-10] π Add Kimi Delta Attention (KDA) implementation to `fla` ([paper](https://arxiv.org/abs/2510.26692)). - [2025-09] π² Add DeltaFormer implementation to `fla` ([paper](https://arxiv.org/abs/2505.19488v1)). - [2025-09] π» Thrilled to announce that [GDN](fla/ops/gated_delta_rule) has been integrated into Qwen3-Next. Check out their [blog post](https://qwen.ai/blog?id=4074cca80393150c248e508aa62983f9cb7d27cd&from=research.latest-advancements-list) for more info! - [2025-08] π² Add Log-Linear Attention implementation to `fla` ([paper](https://arxiv.org/abs/2506.04761)). - [2025-08] π Add MoM implementation to `fla` ([paper](https://arxiv.org/abs/2502.13685)). <details> <summary>Older news</summary> - [2025-07] π³ Add MLA implementation to `fla` ([paper](https://arxiv.org/abs/2405.04434)). - [2025-07] π£οΈ Add PaTH Attention implementation to `fla` ([paper](https://arxiv.org/abs/2505.16381)). - [2025-06] π Add MesaNet implementation to `fla` ([paper](https://arxiv.org/abs/2506.05233)). - [2025-06] π Add Comba implementation to `fla` ([paper](https://arxiv.org/abs/2506.02475)). - [2025-05] π Add Rodimus* implementation to `fla` ([paper](https://arxiv.org/abs/2410.06577)). - [2025-04] π Add DeltaProduct implementation to `fla` ([paper](https://arxiv.org
Excerpt of 42,971 characters
Read on GitHubYu Zhang Β· Moonshot AI Β· China
1.2k
351
Zhiyuan Li
277
sunyi0505
13
ChunyuWei Β· Hangzhou Β· China
12
Eric Alcaide Β· Switzerland
10
9
9
8
L'UniversitΓ© Tsinghua
7
6
Janna Β· United States
6
Ridger Zhu Β· University of California, Santa Cruz
6
Anton Vlasjuk
5
morluto
5
5
5
4
4
Tarushii Goel Β· @thinking-machines-lab
4
Would you bet a product on this? Bounded 0β100 and slow moving.
matched fp:4821f3efaef355cd, topic:large-language-models
matched fp:4821f3efaef355cd, topic:natural-language-processing