Top AI Repos โ open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
๐ This is a repository for organizing papers, codes and other resources related to Visual Reinforcement Learning.
| Date | Stars |
|---|---|
| 2026-07-31 | 452 |
| 2026-08-02 | 452 |
| 2026-08-03 | 452 |
| 2026-08-04 | 453 |
| 2026-08-06 | 453 |
Today
โ stars today
This week
โ stars this week
This month
โ stars this month
Momentum
0.0
growth rate 0.00%/day
# Awesome-RL-for-Multimodal-Foundation-Models
<div align="center">
<img src="assets/logo.png" alt="Logo" width="450">
<p align="center">
This repository accompanies our survey paper: <br><a href="https://arxiv.org/abs/2508.08189"><strong>Reinforcement Learning for Multimodal Foundation Models: A Survey</strong></a>
</p>
</div>
<div align="center">
<a href="https://arxiv.org/abs/2508.08189">
<img src="https://img.shields.io/badge/paper-A42C25?style=for-the-badge&logo=arxiv&logoColor=white" alt="Paper">
</a>
<a href="https://github.com/weijiawu/Awesome-RL-for-Multimodal-Foundation-Models">
<img src="https://img.shields.io/badge/Visual_RL-000000?style=for-the-badge&logo=github&logoColor=000&logoColor=white" alt="Github">
</a>
<a href="https://huggingface.co/papers/2508.08189">
<img src="https://img.shields.io/badge/HuggingFace-fcd022?style=for-the-badge&logo=huggingface&logoColor=000" alt="Hugging Face Collection">
</a>
<a href="https://x.com/weijiawu7/status/1955159048934088713">
<img src="https://img.shields.io/badge/Twitter-%23000000.svg?style=for-the-badge&logo=twitter&logoColor=white" alt="Twitter">
</a>
</div>
<!-- ๐ This is a repository for organizing papers, codes and other resources related to Visual Reinforcement Learning. -->
---
## ๐ News
- **[2026-01-28]** To better define the scope of the claim, we have renamed the title to "Reinforcement Learning for Multimodal Foundation Models: A Survey".
- **[2025-08-13]** We have released ["Reinforcement Learning for Large Model: A Survey"](https://arxiv.org/abs/2508.08189), the **first comprehensive survey** dedicated to the emerging paradigm of "RL for Large Model".
- **[2025-08-13]** We reorganized the repository and aligned the classifications in the survey.
- **[2025-06-08]** We created this repository to maintain a paper list on Awesome-Visual-Reinforcement-Learning. **<span style="color:red">Everyone is welcome to push and update related work!</span>**
---
#### :thinking: What is Reinforcement Learning of Multimodal Foundation Models?
**Reinforcement Learning for Multimodal Foundation Models** enables agents to learn decision-making policies directly from visual observations (e.g., images or videos), rather than structured state inputs.
It lies at the intersection of reinforcement learning and computer vision, with applications in robotics, embodied AI, games, and interactive environments.
#### ๐ Project Description
Awesome-Visual-Reinforcement-Learning is a curated list of papers, libraries, and resources on learning control policies from visual input.
It aims to help researchers and practitioners navigate the fast-evolving Visual RL landscape โ from perception and representation learning to policy learning and real-world applications.
<p align="center">
<img src="assets/timeline.png" alt="TAX" style="display: block; margin: 0 auto;" width="700px" />
</p>
We structure this collection along a trajectory of visual RL. This chart ugroups existing work by high-level domain (MLLMs, visual generation, unified models, and vision-language action agents) and then by finer-grained tasks, illustrating representative papers for each branch.:
<p align="center">
<img src="assets/taxonomy.png" alt="TAX" style="display: block; margin: 0 auto;" width="800px" />
</p>
## ๐ Table of Contents <!-- omit in toc -->
Libraries and tools
- [Benchmarks environments and datasets with Visual RL](#benchmarks-environments-and-datasets-with-visual-rl)
- [Multi-Modal Large Language Models with RL](#multi-modal-large-language-models-with-rl)
- [Visual Generation with RL](#visual-generation-with-rl)
- [RL for Unified Model](#rl-for-unified-model)
- [Vision Language Action Models with RL](#vision-language-action-models-with-rl)
- [Others](#others)
### Benchmarks environments and datasets with Visual RL
#### MLLM
+ [MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning](httpExcerpt of 89,946 characters
Read on GitHubWould you bet a product on this? Bounded 0โ100 and slow moving.
matched fp:af7a028ff412d25a, llm:Repository description: 'organizing papers, codes and other resources related to Visual Reinforcement Learning.' (Awesome list for Visual RL / Multimodal Foundation Models)
matched fp:af7a028ff412d25a, llm:Repository description: 'organizing papers, codes and other resources related to Visual Reinforcement Learning.' (Awesome list for Visual RL / Multimodal Foundation Models)
matched fp:af7a028ff412d25a, llm:Repository description: 'organizing papers, codes and other resources related to Visual Reinforcement Learning.' (Awesome list for Visual RL / Multimodal Foundation Models)