QwenLM/Qwen3-VL
quality grade C, 63 out of 100Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
- stars
- 20k
- stars gained this week
- +40this week
- forks, open issues and contributors
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Vision-language models, document understanding and any-to-any architectures.
Signals: multimodal, vision-language-model, vlm, clip, multimodal-llm, visual-question-answering, document-understanding, image-captioning
340 results
Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
Paper list about multimodal and large language models, only used to record papers I read in the daily arxiv for personal needs.
GPT4Tools is an intelligent system that can automatically decide, control, and utilize different visual foundation models, allowing the user to interact with images during a conversation.
🛰️ Official repository of paper "RemoteCLIP: A Vision Language Foundation Model for Remote Sensing" (IEEE TGRS)
BiomedGPT: A Generalist Vision-Language Foundation Model for Diverse Biomedical Tasks
Vision-Language Pathology Foundation Model - Nature Medicine
Frontier Multimodal Foundation Models for Image and Video Understanding
Transformers 3rd Edition
[T-IV] This repository collects research papers of large Vision Language Models in Autonomous driving and Intelligent Transportation System. The repository will be continuously updated to track the latest update.
Official repo for "More Thought, Less Accuracy? On the Dual Nature of Reasoning in Vision-Language Models" (ICLR 2026)
[NeurIPS 2025]《SD-VLM: Spatial Measuring and Understanding with Depth-encoded Vision Language Models》
This repository collects and organises state‑of‑the‑art papers on spatial reasoning for Multimodal Vision–Language Models (MVLMs).
[ECCV2024] 🐙Octopus, an embodied vision-language model trained with RLEF, emerging superior in embodied visual planning and programming.
Harnessing 1.4M GPT4V-synthesized Data for A Lite Vision-Language Model
[CVPR 2024 Highlight] Mitigating Object Hallucinations in Large Vision-Language Models through Visual Contrastive Decoding
[ICLR'25] MMed-RAG: Versatile Multimodal RAG System for Medical Vision Language Models
The official repo for "SpatialBot: Precise Spatial Understanding with Vision Language Models.
Benchmark environment for evaluating vision-language models (VLMs) on popular video games!
VisualGPT, CVPR 2022 Proceeding, GPT as a decoder for vision-language models
Qianfan-VL: Domain-Enhanced Universal Vision-Language Models
[NeurIPS 2025] Efficient Reasoning Vision Language Models
Experiments and data for the paper "When and why vision-language models behave like bags-of-words, and what to do about it?" Oral @ ICLR 2023
[ICLR'25] Official code for the paper 'MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs'
The code for "TokenPacker: Efficient Visual Projector for Multimodal LLM", IJCV2025
24,538 repositories in the index in total.