deepseek-ai/DeepSeek-VL
quality grade C, 55 out of 100DeepSeek-VL: Towards Real-World Vision-Language Understanding
- stars
- 4.1k
- stars gained this week
- —this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Vision-language models, document understanding and any-to-any architectures.
Signals: multimodal, vision-language-model, vlm, clip, multimodal-llm, visual-question-answering, document-understanding, image-captioning
344 results
DeepSeek-VL: Towards Real-World Vision-Language Understanding
The official repo of Qwen-VL (通义千问-VL) chat & pretrained large vision language model proposed by Alibaba Cloud.
👀「大模型」2小时从0训练65M参数的视觉多模态VLM!Train a 65M-parameter VLM from scratch in just 2h!
[NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
Implementation of CVPR 2023 paper "Prompting Large Language Models with Answer Heuristics for Knowledge-based Visual Question Answering".
MathVista: data, code, and evaluation for Mathematical Reasoning in Visual Contexts
Deep Modular Co-Attention Networks for Visual Question Answering
Knowledge Graphs Meet Multi-Modal Learning: A Comprehensive Survey
X-modaler is a versatile and high-performance codebase for cross-modal analytics(e.g., image captioning, video captioning, vision-language pre-training, visual question answering, visual commonsense reasoning, and cross-modal retrieval).
Implementation of 🦩 Flamingo, state-of-the-art few-shot visual question answering attention net out of Deepmind, in Pytorch
Bottom-up attention model for image captioning and VQA, based on Faster R-CNN and Visual Genome
Official repository of OFA (ICML 2022). Paper: OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework
[CVPR 2026] LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
[CVPR 2026] Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens
2D to 3D CAD Conversion Using VLM
VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs
🏭 Mega Scale Multimodal DataPipeline for SOTA Foundation Models
[ICML 2026 Oral] Holi-Spatial: Evolving Video Streams into Holistic 3D Spatial Intelligence
EVE Series: Encoder-Free Vision-Language Models from BAAI
Awesome Multimodal Modeling [Covers MLLM, UMM, and NMM]
Flame is an open-source multimodal AI system designed to translate UI design mockups into high-quality React code. It leverages vision-language modeling, automated data synthesis, and structured training workflows to bridge the gap between design and front-end development.
ScreenAgent: A Computer Control Agent Driven by Visual Language Large Model (IJCAI-24)
This repo contains the code for "VLM2Vec / MMEB" [ICLR 2025], "VLM2Vec-V2 / MMEB-V2" [TMLR 2026], and "MMEB-V3" [COLM 2026]
24,523 repositories in the index in total.