Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
| Date | Stars |
|---|---|
| 2026-07-31 | 531 |
| 2026-08-06 | 531 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
## <img src="./assets/vlnr1_logo.png" alt="Icon" width="70" height="70"> VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning
<div style="text-align: center;">
<p class="title is-5 mt-2 authors">
<a href="https://scholar.google.com/citations?user=kwVLpo8AAAAJ&hl=en/" target="_blank">Zhangyang Qi</a><sup>1,2*</sup>,
<a href="https://github.com/rookiexiong7" target="_blank">Zhixiong Zhang</a><sup>2</sup>,
<a href="https://i.cs.hku.hk/~yzyu/publication/" target="_blank">Yizhou Yu</a><sup>1</sup>,
<a href="https://myownskyw7.github.io/" target="_blank">Jiaqi Wang</a><sup>2✉</sup>,
<a href="https://hszhao.github.io/" target="_blank">Hengshuang Zhao</a><sup>1✉</sup>
</p>
</div>
<div style="text-align: center;">
<!-- contribution -->
<p class="subtitle is-5" style="font-size: 1.0em; text-align: center;">
<sup>*</sup> Equation Contribution,
<sup>✉</sup> Corresponding Authors,
</p>
</div>
<div style="text-align: center;">
<!-- affiliations -->
<p class="subtitle is-5" style="font-size: 1.0em; text-align: center;">
<sup>1</sup> The University of Hong Kong,
<sup>2</sup> Shanghai AI Laboratory,
</p>
</div>
<p align="center">
<a href="https://arxiv.org/abs/2506.17221" target='_**blank**'>
<img src="https://img.shields.io/badge/arXiv-2506.17221📖-bron?">
</a>
<a href="https://vlnr1.github.io/" target='_blank'>
<img src="https://img.shields.io/badge/Project%20page-🚀-yellow">
</a>
<a href="https://huggingface.co/datasets/alexzyqi/VLN-R1-datasets/" target='_blank'>
<img src="https://img.shields.io/badge/Huggingface%20Datasets-🤗-blue">
</a>
<a href="https://x.com/Qi_Zhangyang" target='_blank'>
<img src="https://img.shields.io/twitter/follow/Qi_Zhangyang">
</a>
</p>
## <img src="./assets/gptscene_logo.png" alt="Icon" width="80" height="40"> [ICLR 2026] GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
<div style="text-align: center;">
<p class="title is-5 mt-2 authors">
<a href="https://scholar.google.com/citations?user=kwVLpo8AAAAJ&hl=en/" target="_blank">Zhangyang Qi</a><sup>1,2*</sup>,
<a href="https://github.com/rookiexiong7" target="_blank">Zhixiong Zhang</a><sup>2*</sup>,
<a href="https://github.com/Aleafy" target="_blank">Ye Fang</a><sup>2</sup>,
<a href="https://myownskyw7.github.io/" target="_blank">Jiaqi Wang</a><sup>2✉</sup>,
<a href="https://hszhao.github.io/" target="_blank">Hengshuang Zhao</a><sup>1✉</sup>
</p>
</div>
<div style="text-align: center;">
<!-- contribution -->
<p class="subtitle is-5" style="font-size: 1.0em; text-align: center;">
<sup>*</sup> Equation Contribution,
<sup>✉</sup> Corresponding Authors,
</p>
</div>
<div style="text-align: center;">
<!-- affiliations -->
<p class="subtitle is-5" style="font-size: 1.0em; text-align: center;">
<sup>1</sup> The University of Hong Kong,
<sup>2</sup> Shanghai AI Laboratory,
</p>
</div>
<p align="center">
<a href="https://arxiv.org/abs/2501.01428" target='_**blank**'>
<img src="https://img.shields.io/badge/arXiv-2501.01428📖-bron?">
</a>
<a href="https://gpt4scene.github.io/" target='_blank'>
<img src="https://img.shields.io/badge/Project%20page-🚀-yellow">
</a>
<a href="https://huggingface.co/alexzyqi/GPT4Scene-qwen2vl_full_sft_mark_32_3D_img512" target='_blank'>
<img src="https://img.shields.io/badge/Huggingface%20Models-🤗-blue">
</a>
<a href="https://x.com/Qi_Zhangyang" target='_blank'>
<img src="https://img.shields.io/twitter/follow/Qi_Zhangyang">
</a>
</p>
## 🔥 News
[2026/01/26] GPT4Scene is accepted to ICLR 2026!
[2025/07/06] We have released the **[training data](https://huggingface.co/datasets/alexzyqi/VLN-R1-datasets/)**, **[tokenizers](https://huggingface.co/datasets/alexzyqi/GPT4Scene_VLN-R1_tokenizers/)** and **[data generation code](https://github.com/Qi-ZhangyExcerpt of 8,959 characters
Read on GitHubQi Zhangyang (Alex) · The University of Hong Kong (HKU ) · Hong Kong
15
3
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:e1d81d64c83682db, desc:vision-language