jingyi0000/R1-VL
quality grade D, 47 out of 100R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization
- stars
- 352
- stars gained this week
- —
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Vision-language models, document understanding and any-to-any architectures.
Signals: multimodal, vision-language-model, vlm, clip, multimodal-llm, visual-question-answering, document-understanding, image-captioning
344 results
R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization
Recipes for learning, fine-tuning, and adapting ColPali to your multimodal RAG use cases. 👨🏻🍳
[ICLR 2026 Oral (top 1.2%)] Official implementation of DepthLM
Official repository for "VideoPrism: A Foundational Visual Encoder for Video Understanding" (ICML 2024)
MOSS-VL is the core multimodal model series within the OpenMOSS ecosystem, dedicated to visual understanding.
[CVPR 2024 Highlight] OPERA: Alleviating Hallucination in Multi-Modal Large Language Models via Over-Trust Penalty and Retrospection-Allocation
InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy
[ECCV 2026] Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding
[CVPR 2025 Highlight] Official code for "Olympus: A Universal Task Router for Computer Vision Tasks"
An open-source implementation for training LLaVA-NeXT.
Official repository for VisionZip (CVPR 2025)
Bridging Large Vision-Language Models and End-to-End Autonomous Driving
LLaVA-Mini is a unified large multimodal model (LMM) that can support the understanding of images, high-resolution images, and videos in an efficient manner.
Evaluating text-to-image/video/3D models with VQAScore
A curated list of awesome prompt/adapter learning methods for vision-language models like CLIP.
[CVPR 2024 🔥] Grounding Large Multimodal Model (GLaMM), the first-of-its-kind model capable of generating natural language responses that are seamlessly integrated with object segmentation masks.
Fully Open Framework for Democratized Multimodal Training
Awesome Unified Multimodal Models
The implementation of "Prismer: A Vision-Language Model with Multi-Task Experts".
[ICCV 2025] Implementation for Describe Anything: Detailed Localized Image and Video Captioning
Get clean data from tricky documents, powered by vision-language models ⚡
Seed1.5-VL, a vision-language foundation model designed to advance general-purpose multimodal understanding and reasoning, achieving state-of-the-art performance on 38 out of 60 public benchmarks.
A curated list of awesome LLM/VLM/VLA/World Model for Autonomous Driving(LLM4AD) resources (continually updated)
Official repo for "Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models"
24,523 repositories in the index in total.