Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
[CVPR 2026] LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
| Date | Stars |
|---|---|
| 2026-07-24 | 255 |
| 2026-07-25 | 255 |
| 2026-07-28 | 257 |
| 2026-07-30 | 258 |
| 2026-08-06 | 258 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling <div align="center"> [](https://arxiv.org/abs/2511.20785) [](https://evolvinglmms-lab.github.io/LongVT/) [](https://github.com/EvolvingLMMs-Lab/LongVT) [](https://huggingface.co/datasets/longvideotool/LongVT-Parquet) [](https://huggingface.co/collections/lmms-lab/longvt) [](https://huggingface.co/spaces/longvideotool/LongVT-Demo) [](https://www.lmms-lab.com/posts/longvt/) [](https://huggingface.co/papers/2511.20785) </div> ## 🎉 News - **[2026-05-20]**: We released [**ParaVT**](https://github.com/EvolvingLMMs-Lab/ParaVT), our follow-up work that extends LongVT from sequential to parallel video tool calling, post-trained with **PARA-GRPO** to tame the *Tool Prior Paradox* (Format Fragility + Tool Necessity Gap) of native-RL agentic video reasoning. - **[2026-03-09]**: We fixed SFT parquet schema issue ([#14](https://github.com/EvolvingLMMs-Lab/LongVT/issues/14)) and released deduplicated [VideoSIAH-Eval](https://huggingface.co/datasets/longvideotool/VideoSIAH-Eval). - **[2026-02-21]**: LongVT was accepted by 🔥 **CVPR 2026**! - **[2026-01-25]**: We are invited to **AAAI Talk**! Check out the [slides](https://docs.google.com/presentation/d/1Xm0tH28hdZKBLB7d5LCNrFJNJQtd6kasKPov15FH4FE/edit?usp=sharing). - **[2025-12-10]**: We are invited to **BAAI Talk**! Check out the [slides and recording](https://event.baai.ac.cn/activities/983). - **[2025-12-07]**: We won the 🏆 **[AI Paper of the Day](https://huggingface.co/collections/vladbogo/ai-paper-of-the-day)** (on Dec 02, 2025), **Top \#3 Weekly Paper** (by Dec 07, 2025), and **Top \#5 Monthly Paper** (in December, 2025) on Hugging Face! Check out the [LongVT Paper Page](https://huggingface.co/papers/2511.20785). - **[2025-11-28]**: We released all of our codes, data, and model checkpoints! Check out the [LongVT Collection](https://huggingface.co/collections/lmms-lab/longvt). ## Table of Contents - [Overview](#overview) - [Installation](#installation) - [SFT Training](#1-sft-training) - [RL Training](#2-rl-training) - [RFT Training](#3-rft-training) - [Evaluation](#4-evaluation) - [Data Pipeline](#5-data-pipeline) - [Getting Started](#getting-started) - [Data Preparation](#data-preparation) - [SFT Training](#sft-training) - [RL Training](#rl-training) - [RFT Training](#rft-training) - [Evaluation](#evaluation) - [LLM Judge Setup](#llm-judge-setup) - [Data Pipeline](#data-pipeline) - [Single Sample Inference](#single-sample-inference) - [FAQ](#faq) - [Citation](#citation) - [Acknowledgements](#acknowledgements) - [Star History](#-star-history) ## Overview <div align="center"> <img src="assets/teaser.png" alt="teaser" width="1000"/> </div> Large multimodal models (LMMs) have shown great potential for video reasoning with textual Chain-of-Thought. However, they remain vulnerable to hallucinations, especially when processing long-form videos where evidence is sparse and temporally dispersed. Inspired by how humans comprehend long videos—by first skimming globally and then examining relevant clips for details—we introduce **LongVT**, an end-to-end agentic framework that enables "Thinking with **Long V**
Excerpt of 27,580 characters
Read on GitHubWould you bet a product on this? Bounded 0–100 and slow moving.
matched fp:55310aeaeed671cd, topic:multimodal, topic:vlm, readme:multimodal