Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
[ICLR 2026] On the Generalization of SFT: A Reinforcement Learning Perspective with Reward Rectification.
| Date | Stars |
|---|---|
| 2026-07-31 | 614 |
| 2026-08-01 | 794 |
| 2026-08-02 | 1084 |
| 2026-08-03 | 1085 |
| 2026-08-05 | 1095 |
| 2026-08-06 | 1095 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
<div align="center">
# *On the Generalization of SFT*: <br>A Reinforcement Learning Perspective with <br>Reward Rectification
<a href="http://arxiv.org/abs/2508.05629" target="_blank">
<img alt="arXiv" src="https://img.shields.io/badge/arXiv-DFT-red?logo=arxiv" height="25" />
</a>
<a href="https://huggingface.co/collections/Liang0223/dft-6892da5e421a56a8deb48c9f" target="_blank">
<img alt="Hugging Face Models" src="https://img.shields.io/badge/%F0%9F%A4%97%20_Huggingface-Models-ffc107?color=ffc107&logoColor=white" height="25" />
</a>
<div style="font-family: charter; text-align: center; margin: 0 auto;">
<a href="https://yongliang-wu.github.io/" class="author-link" target="_blank">Yongliang Wu*</a>  
<a href="https://scholar.google.com/citations?user=dHBNmSkAAAAJ" class="author-link" target="_blank">Yizhou Zhou*</a>  
<a href="https://scholar.google.com/citations?user=IH2wK1cAAAAJ" class="author-link" target="_blank">Zhou Ziheng</a>  
<a href="https://github.com/ForJadeForest" class="author-link" target="_blank">Yingzhe Peng</a>  
<br>
<a href="https://scholar.google.com/citations?user=fdwhd9gAAAAJ" class="author-link" target="_blank">Xinyu Ye</a>  
<a href="https://joyhuyy1412.github.io/" class="author-link" target="_blank">Xinting Hu</a>  
<a href="https://vitozhu04.github.io/" class="author-link" target="_blank">Wenbo Zhu</a>  
<a href="http://luqi.info/" class="author-link" target="_blank">Lu Qi</a>  
<a href="https://faculty.ucmerced.edu/mhyang/" class="author-link" target="_blank">Ming-Hsuan Yang</a>  
<a href="https://yxpalmweb.github.io/" class="author-link" target="_blank">Xu Yang</a>  
</div>
<br>
</div>
## 🌟 Thanks for the Feedback of Community
We are grateful for the many thoughtful comments and feedback from the community regarding DFT, ranging from discussions of related ideas to reports of its application in different scenarios. We have heard of both successes and failures when applying DFT, for instance in literary or financial tasks.
Here, we would like to clarify that we do not claim DFT can replace SFT in all cases, as noted in our limitations section:
> *“While our experiments demonstrate substantial gains from DFT on mathematical reasoning benchmarks, this evaluation is confined to math-focused and code-focused (will be released in next version) datasets and models up to 7 billion parameters.”*
---
Nonetheless, these less successful cases, as well as community discussions on platforms such as Zhihu or Xiao Hong Shu about the intuitive principles behind DFT, together with our own experimental experience, have prompted us to think more deeply about the conditions under which DFT works well, and why it may be less effective in other contexts.
All this feedback reminds us of a remark by computing pioneer Richard Hamming in *The Art of Doing Science and Engineering: Learning to Learn* (p.27), which we have slightly adapted:
> *“Almost everyone who opens up a new field does not really understand it the way the followers—or the critics—do.”*
---
We hope this work can contribute to renewed interest in exploring the interplay between SFT and RL, and in better understanding the factors that underlie both the successes and the limitations of methods like DFT. Looking ahead, we also welcome researchers who are interested in our work to improve DFT in some of the currently unsuccessful cases, or in leveraging the ideas to uncover other connections between RL algorithms and SFT, ultimately aiming to achieve RL-like benefits at the cost of SFT across a broader range of settings.
## 📰 News
* **\[2025.08.08]** We have released the training scripts, evaluation scripts, and model checkpoints.
## Abstract
We present a simple yet thExcerpt of 9,909 characters
Read on GitHub30
2
1
1
1
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:d4b9d6923f436691, desc:reinforcement learning