Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
A Survey of Reinforcement Learning for Large Reasoning Models
| Date | Stars |
|---|---|
| 2026-07-24 | 2468 |
| 2026-07-25 | 2469 |
| 2026-07-28 | 2469 |
| 2026-07-30 | 2469 |
| 2026-07-31 | 2472 |
| 2026-08-02 | 2473 |
| 2026-08-03 | 2475 |
| 2026-08-06 | 2475 |
Today
— stars today
This week
+6 stars this week
This month
— stars this month
Momentum
6.0
growth rate 0.24%/day
<div align="center"> <img src="figs/survey_logo.png" style="width: 70%;"/> ## A Survey of Reinforcement Learning for Large Reasoning Models [](https://github.com/sindresorhus/awesome) [](https://arxiv.org/abs/2509.08827) [](https://github.com/TsinghuaC3I/Awesome-RL-Reasoning-Recipes) [](https://huggingface.co/papers/2509.08827) [](https://x.com/OkhayIea/status/1965989894163235111) </div> > We welcome everyone to open an issue for any related work we haven’t discussed, and we’ll try to address it in the next release! ## 🎉 News - **[2025-11-05]** 🔥 Excited to release our paper list about **Memory for Agents**, covering breakthroughs in Context Management and Learning from Experience powering self-improving AI agents. Check it out: [GitHub](https://github.com/TsinghuaC3I/Awesome-Memory-for-Agents) - **[2025-10]** 🎉 Honored to give talks at [BAAI](https://event.baai.ac.cn/activities/961), [Qingke Talk](https://qingkeai.online/archives/0h3Cm8Bi) and Tencent Wiztalk! Here are the [slides]([email protected]). - **[2025-09-18]** 🎉 We update the full list of papers in the category structure of the survey! - **[2025-09-12]** 🎉 Our survey was ranked **#1 Paper of the Day** on 🤗 [Hugging Face Daily Papers](https://huggingface.co/papers/2509.08827)! - **[2025-09-11]** 🔥 Excited to release our **RL for LRMs Survey**! We’ll be updating the full list of papers in with a new category structure soon. Check it out: [Paper](https://huggingface.co/papers/2509.08827). - **[2025-08-15]** 🔥 Introducing **SSRL**: an investigation for Agentic Search RL without reliance on external search engine. Check it out: [GitHub](https://github.com/TsinghuaC3I/SSRL) and [Paper](https://arxiv.org/abs/2508.10874). - **[2025-05-27]** 🔥 Introducing **MARTI**: A Framework for LLM-based Multi-Agent Reinforced Training and Inference. Check it out: [Github](https://github.com/TsinghuaC3I/MARTI). - **[2025-04-23]** 🔥 Introducing **TTRL**: an open-source solution for online RL on data without ground-truth labels, especially test data. Check it out: [Github](https://github.com/PRIME-RL/TTRL) and [Paper](https://arxiv.org/abs/2504.16084). - **[2025-03-20]** 🔥 We are excited to introduce collection of papers and projects on RL for reasoning models! ## 🎈 Citation If you find this survey helpful, please cite our work: ```bibtex @article{zhang2025survey, title={A survey of reinforcement learning for large reasoning models}, author={Zhang, Kaiyan and Zuo, Yuxin and He, Bingxiang and Sun, Youbang and Liu, Runze and Jiang, Che and Fan, Yuchen and Tian, Kai and Jia, Guoli and Li, Pengfei and others}, journal={arXiv preprint arXiv:2509.08827}, year={2025} } ``` ## 📖 Contents - [A Survey of Reinforcement Learning for Large Reasoning Models](#a-survey-of-reinforcement-learning-for-large-reasoning-models) - [🎉 News](#-news) - [🎈 Citation](#-citation) - [📖 Contents](#-contents) - [🗺️ Overview](#️-overview) - [📄 Paper List](#-paper-list) - [Frontier Models](#frontier-models) - [Reward Design](#reward-design) - [Generative Rewards](#generative-rewards) - [Dense Rewards](#dense-rewards) - [Unsupervised Rewards](#unsupervised-rewards) - [Rewards Shaping](#rewards-shaping) - [Policy Optimization](#policy-optimization) - [Policy Gradient Objective](#policy-gradient-objective) - [Critic-based Algorithms](#critic-based-algorithms) - [Critic-Free Algorithms](#critic-free-algorithms) - [Off-policy O
Excerpt of 206,913 characters
Read on GitHubKaiyan Zhang · @FrontisAI · China
106
Yuxin Zuo
28
18
13
6
zclzc · @google-deepmind · Singapore
4
chrisliu298 · University of California, Santa Cruz · United States
4
4
3
3
Jiacheng Lin · University of Illinois Urbana-Champaign · United States
3
Zhiwei He · Shanghai Jiao Tong University · China
2
2
2
2
George
2
1
1
Yi-Fan Zhang · State Key Laboratory of Pattern Recognition
1
Wujiang Xu · Meta, Rutgers University · United States
1
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:0ddf3564c27c5b8a, topic:rl, desc:reinforcement learning, readme:reinforcement learning
matched fp:0ddf3564c27c5b8a, topic:llm
matched fp:0ddf3564c27c5b8a, topic:awesome-list, readme:paper list