Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Official implementation of Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning
| Date | Stars |
|---|---|
| 2026-07-31 | 276 |
| 2026-08-02 | 276 |
| 2026-08-03 | 276 |
| 2026-08-04 | 276 |
| 2026-08-06 | 276 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
<div align="center">
<h1 align="center"> Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning
</h1>
[UnifiedReward](https://github.com/CodeGoat24/UnifiedReward) Team
<a href="https://arxiv.org/pdf/2508.20751">
<img src='https://img.shields.io/badge/arXiv-Pref GRPO-blue' alt='Paper PDF'></a>
<a href="https://codegoat24.github.io/UnifiedReward/Pref-GRPO">
<img src='https://img.shields.io/badge/Project-Website-orange' alt='Project Page'></a>
<a href="https://huggingface.co/collections/CodeGoat24/unifiedreward-flex">
<img src='https://img.shields.io/badge/Huggingface-UnifiedReward Flex-yellow' alt='Project Page'></a>
<a href="https://huggingface.co/collections/CodeGoat24/pref-grpo-and-unigenbench">
<img src='https://img.shields.io/badge/Huggingface-Model-yellow' alt='Project Page'></a>
<a href="https://github.com/CodeGoat24/UniGenBench">
<img src='https://img.shields.io/badge/Benchmark-UniGenBench-green' alt='Project Page'></a>
[-brown)](https://huggingface.co/spaces/CodeGoat24/UniGenBench_Leaderboard)
[-red)](https://huggingface.co/spaces/CodeGoat24/UniGenBench_Leaderboard_Chinese)
[-orange)](https://huggingface.co/spaces/CodeGoat24/UniGenBench_Leaderboard_English_Long)
[-pink)](https://huggingface.co/spaces/CodeGoat24/UniGenBench_Leaderboard_Chinese_Long)
</div>
## 🔥 News
Please leave us a star if you find this work helpful.
- [2026/02] Support **Z-Image**, **FLUX.1-Kontext-dev**, **FLUX.2-Klein (T2I/I2I)**, **Qwen-Image-Edit** and **Wan2.2**.
- [2026/02] Support [UnifiedReward-Flex](https://codegoat24.github.io/UnifiedReward/flex)-based Pref-GRPO for both image and video generation.
- [2026/01] **Tongyi Lab** improves Pref-GRPO on open-ended agents in [ArenaRL: Scaling RL for Open-Ended Agents via Tournament-based Relative Ranking](https://arxiv.org/pdf/2601.06487). Thanks to all contributors!
<details>
<summary><strong>More News</strong></summary>
- [2025/11] Support **Qwen-Image**, **Wan2.1** and **FLUX.1-dev**.
- [2025/11] **Nano Banana Pro**, **FLUX.2-dev** and **Z-Image** are added to all Leaderboards.
- [2025/10] **Alibaba Group** proves the effectiveness of Pref-GRPO on aligning LLMs in [Taming the Judge: Deconflicting AI Feedback for Stable Reinforcement Learning](https://arxiv.org/pdf/2510.15514). Thanks to all contributors!
- [2025/9] **Seedream-4.0**, **GPT-4o**, **Imagen-4-Ultra**, **Nano Banana**, **Lumina-DiMOO**, **OneCAT**, **Echo-4o**, **OmniGen2**, and **Infinity** are added to all Leaderboards.
- [2025/8] Release [Leaderboard (**English**)](https://huggingface.co/spaces/CodeGoat24/UniGenBench_Leaderboard), [Leaderboard (**English Long**)](https://huggingface.co/spaces/CodeGoat24/UniGenBench_Leaderboard_English_Long), [Leaderboard (**Chinese Long**)](https://huggingface.co/spaces/CodeGoat24/UniGenBench_Leaderboard_Chinese_Long) and [Leaderboard (**Chinese**)](https://huggingface.co/spaces/CodeGoat24/UniGenBench_Leaderboard_Chinese).
</details>


## 🔧 Environment Setup
1. Clone this repository and navigate to the folder:
```bash
git clone https://github.com/CodeGoat24/Pref-GRPO.git
cd Pref-GRPO
```
2. Install the training package:
```bash
conda create -n PrefGRPO python=3.12
conda activate PrefGRPO
bash env_setup.sh fastvideo
git clone https://github.com/mlfoundations/open_clip
cd open_clip
pip install -e .
cd ..
```
3. Install vLLM (for UnifiedReward-based rewards)
```bash
conda create -n vllm
conda actExcerpt of 16,778 characters
Read on GitHubWould you bet a product on this? Bounded 0–100 and slow moving.
matched fp:7f52345735bd1dfd, name:grpo, desc:grpo