Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Official repository of 'Visual-RFT: Visual Reinforcement Fine-Tuning' & 'Visual-ARFT: Visual Agentic Reinforcement Fine-Tuning'’
| Date | Stars |
|---|---|
| 2026-07-31 | 2263 |
| 2026-08-06 | 2265 |
Today
+2 stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
<p align="center">
<!-- <h1 align="center"><img src="assets/logo.png" width="256"></h1> -->
<h1 align="center">Visual-RFT: Visual Reinforcement Fine-Tuning</h1>
<p align="center">
<a href="https://github.com/Liuziyu77"><strong>Ziyu Liu*</strong></a>
·
<a href="https://github.com/SunzeY"><strong>Zeyi Sun*</strong></a>
·
<a href="https://yuhangzang.github.io/"><strong>Yuhang Zang</strong></a>
·
<a href="https://lightdxy.github.io/"><strong>Xiaoyi Dong</strong></a>
·
<a href="https://scholar.google.com/citations?user=sJkqsqkAAAAJ"><strong>Yuhang Cao</strong></a>
·
<a href="https://kennymckormick.github.io/"><strong>Haodong Duan</strong></a>
·
<a href="http://dahua.site/"><strong>Dahua Lin</strong></a>
·
<a href="https://myownskyw7.github.io/"><strong>Jiaqi Wang</strong></a>
</p>
<h2 align="center">Accepted By ICCV 2025!</h2>
<!-- 🏠<a href="https://liuziyu77.github.io/MIA-DPO/">Homepage</a></h3>| -->
📖<a href="https://arxiv.org/abs/2503.01785">Paper</a> |
🤗<a href="https://huggingface.co/collections/laolao77/virft-datasets-67bc271b6f2833eccc0651df">Datasets</a> | 🤗<a href="https://huggingface.co/papers/2503.01785">Daily Paper</a></h3>
<div align="center"></div>
<p align="center">
<p>
🌈We introduce <strong>Visual Reinforcement Fine-tuning (Visual-RFT)</strong>, the first comprehensive adaptation of <strong>Deepseek-R1's RL strategy</strong> to the <strong>multimodal field</strong>. We use the Qwen2-VL-2/7B model as our base model and design a <strong>rule-based verifiable reward</strong>, which is integrated into a <strong>GRPO-based reinforcement fine-tuning framework</strong> to enhance the performance of LVLMs across various visual perception tasks. <strong>ViRFT</strong> extends R1's reasoning capabilities to multiple visual perception tasks, including various detection tasks like <strong>Open Vocabulary Detection, Few-shot Detection, Reasoning Grounding, and Fine-grained Image Classification</strong>.
</p>
<!-- <a href="">
<img src="assets/teaser.png" alt="Logo" width="100%">
</a> -->
<br>
<a href="">
<img src="assets/radar.png" alt="Logo" >
</a>
## 🔥🔥🔥 Visual-RFT: Visual Reinforcement Fine-Tuning
We introduce *Visual Reinforcement Fine-tuning (Visual-RFT)*, the first comprehensive adaptation of Deepseek-R1’s RL strategy to the multimodal field. We use the Qwen2-VL-2/7B model as our base model and design a rule-based verifiable reward, which is integrated into a GRPO-based reinforcement fine-tuning framework to enhance the performance of LVLMs across various visual perception tasks.
📖<a href="https://arxiv.org/abs/2503.01785">Paper</a> | 🤗<a href="https://huggingface.co/collections/laolao77/virft-datasets-67bc271b6f2833eccc0651df">Datasets</a> | 🤗<a href="https://huggingface.co/papers/2503.01785">Daily Paper</a>
## 🔥🔥🔥 Visual-ARFT: Visual Agentic Reinforcement Fine-Tuning
Our new work *Visual Agentic Reinforcement Fine-Tuning (Visual-ARFT)* is designed for enabling flexible and adaptive agentic abilities for Large Vision-Language Models (LVLMs). With Visual-ARFT, open-source LVLMs gain the ability to browse websites for real-time information updates and write code to manipulate and analyze input images through cropping, rotation, and other image processing techniques. We also present a Multi-modal Agentic Tool Bench (MAT) with two settings (MAT-Search and MAT-Coding) designed to evaluate LVLMs’ agentic search and coding abilities.
📖<a href="https://arxiv.org/abs/2505.14246">Paper</a> | 🤗<a href="https://huggingface.co/datasets/laolao77/MAT">Datasets</a> | 🤗<a href="https://huggingface.co/collections/laolao77/visual-arft-682c601d0e35ac6470adfe9f">Models</a>
## 📢 News
- 🚀 [06/26/2025] Our paper **Visual-RFT** is accepted by ICCV 2025!
- 🚀 [05/21/2025] We support both **HuggingFace Dataset** format and **JSON** file format as input datasets for training.
- 🚀 [05/21/2025] We updata the trainer of **VisualExcerpt of 18,985 characters
Read on GitHub137
8
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:108d4bac26eaefbe, desc:fine-tuning, desc:fine tuning