Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
A curated list of reinforcement learning with human feedback resources (continually updated)
| Date | Stars |
|---|---|
| 2026-07-24 | 4416 |
| 2026-07-25 | 4416 |
| 2026-07-28 | 4418 |
| 2026-07-30 | 4418 |
| 2026-07-31 | 4420 |
| 2026-08-02 | 4421 |
| 2026-08-05 | 4422 |
| 2026-08-08 | 4421 |
| 2026-08-10 | 4421 |
| 2026-08-11 | 4421 |
| 2026-08-12 | 4422 |
| 2026-08-15 | 4420 |
| 2026-08-18 | 4422 |
| 2026-08-22 | 4423 |
| 2026-08-27 | 4422 |
| 2026-08-28 | 4421 |
| 2026-08-29 | 4422 |
| 2026-09-02 | 4424 |
| 2026-09-03 | 4425 |
| 2026-09-05 | 4426 |
| 2026-09-06 | 4425 |
| 2026-09-12 | 4427 |
| 2026-09-13 | 4426 |
| 2026-09-14 | 4425 |
| 2026-09-15 | 4426 |
| 2026-09-17 | 4427 |
| 2026-09-18 | 4428 |
| 2026-09-19 | 4427 |
| 2026-09-20 | 4426 |
Today
-1 stars today
This week
— stars this week
This month
+4 stars this month
Momentum
0.0
growth rate 0.00%/day
# Awesome RLHF (RL with Human Feedback)
[](https://github.com/sindresorhus/awesome)    [](https://github.com/opendilab/awesome-RLHF/blob/main/LICENSE)
This is a collection of research papers for **Reinforcement Learning with Human Feedback** (RLHF).
And the repository will be continuously updated to track the frontier of RLHF.
Welcome to follow and star!
## Table of Contents
- [Awesome RLHF (RL with Human Feedback)](#awesome-rlhf-rl-with-human-feedback)
- [Table of Contents](#table-of-contents)
- [Overview of RLHF](#overview-of-rlhf)
- [Detailed Explanation](#detailed-explanation)
- [Papers](#papers)
- [2026](#2026)
- [2025](#2025)
- [2024](#2024)
- [2023](#2023)
- [2022](#2022)
- [2021](#2021)
- [2020 and before](#2020-and-before)
- [Codebases](#codebases)
- [Dataset](#dataset)
- [Blogs](#blogs)
- [Books](#books)
- [Other Language Support](#other-language-support)
- [Contributing](#contributing)
- [License](#license)
## Overview of RLHF
The idea of RLHF is to use methods from reinforcement learning to directly optimize a language model with human feedback. RLHF has enabled language models to begin to align a model trained on a general corpus of text data to that of complex human values.
- RLHF for Large Language Model (LLM)

- RLHF for Video Game (e.g. Atari)

### Detailed Explanation
**(The following section was automatically generated by ChatGPT)**
RLHF typically refers to "Reinforcement Learning with Human Feedback". Reinforcement Learning (RL) is a type of machine learning that involves training an agent to make decisions based on feedback from its environment. In RLHF, the agent also receives feedback from humans in the form of ratings or evaluations of its actions, which can help it learn more quickly and accurately.
RLHF is an active research area in artificial intelligence, with applications in fields such as robotics, gaming, and personalized recommendation systems. It seeks to address the challenges of RL in scenarios where the agent has limited access to feedback from the environment and requires human input to improve its performance.
Reinforcement Learning with Human Feedback (RLHF) is a rapidly developing area of research in artificial intelligence, and there are several advanced techniques that have been developed to improve the performance of RLHF systems. Here are some examples:
- `Inverse Reinforcement Learning (IRL)`: IRL is a technique that allows the agent to learn a reward function from human feedback, rather than relying on pre-defined reward functions. This makes it possible for the agent to learn from more complex feedback signals, such as demonstrations of desired behavior.
- `Apprenticeship Learning`: Apprenticeship learning is a technique that combines IRL with supervised learning to enable the agent to learn from both human feedback and expert demonstrations. This can help the agent learn more quickly and effectively, as it is able to learn from both positive and negative feedback.
- `Interactive Machine Learning (IML)`: IML is a technique that involves active interaction between the agent and the human expert, allowing the expert to provide feedback on the agent's actions in real-time. This can help the agent learn more quickly and efficiently, as it can receive feedback on its actions at each step of the learning process.
- `Human-in-the-Loop Reinforcement Learning (HITLRL)`: HITLRL is aExcerpt of 90,886 characters
Read on GitHubAron751 · AWS
14
Swain · @opendilab · China
10
CloudChen
8
Chen-Hao (Leo) Chang · @Tencent-Hunyuan, @OpenDILab, @vision-x-nyu, HUST · United States
6
6
Wei Xiong
3
Ziyi Zhang · Wuhan University
3
Xu Jingxin
2
South Korea
2
2
2
Stjepan Jureković · Manning Publication · Croatia
2
1
1
1
1
Jia Ruonan
1
1
Feifan Song · Peking University
1
shizhediao · Thinking Machines Lab · United States
1
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:637053c98a2fe394, topic:reinforcement-learning, topic:deep-reinforcement-learning, desc:reinforcement learning
matched fp:637053c98a2fe394, topic:rlhf, name:rlhf, readme:rlhf