A Comprehensive Survey & Curated List of Multimodal Modeling
From Traditional Fusion to Native & Unified Architectures
Overview · Traditional · MLLMs · UMMs · NMMs ·Closed Source Models · Resources
- [2026-06-06] 🚀 We released two new open-source model collections:
- Unified Multimodal Model Zoo (UMM Zoo) – a curated collection of unified multimodal models for understanding, generation, reasoning, and editing.
- Native Multimodal Model Zoo (NMM Zoo) – a curated collection of native multimodal models that natively integrate perception, reasoning, and generation.
- [2026-04-13] ⭐ The repository has already gained over 100 stars in just one day! Thank you all for the incredible support. We will keep updating this list with more cutting-edge models and resources. Your continued stars and PRs are warmly welcomed!
- [2026-04-12] 🎉 We are excited to launch Awesome Multimodal Modeling — a curated reading list organized by architectural paradigms. A comprehensive survey paper is coming soon! Stay tuned.
Browse the list
- Awesome Multimodal Modeling
- 📢 News
- Table of Contents
- About This List
- 1. Introduction & Definitions
- 2. Traditional Multimodal Models
- 3. Multimodal Large Language Models (MLLMs)
- 4. Unified Multimodal Models (UMMs)
- 5. Native Multimodal Models (NMMs)
- 6. Closed-Source Multimodal Models
- 7. Resources
- How to Contribute
- Citation
- License
In this section: At a Glance · Curation Principles
This repository provides a structured, community-maintained survey of multimodal models, covering the full evolutionary arc from early fusion methods to today's natively-trained omni-models. We emphasize precise architectural definitions and classification, especially for the often-conflated categories of Unified Multimodal Models (UMMs) and Native Multimodal Models (NMMs).
Scope: Primary focus on image + text modalities; audio/video/3D are annotated where present. Omni/any-to-any models are marked with Omni.
| Dimension | Coverage |
|---|---|
| Primary scope | Image + text multimodal models, with explicit annotations for video, audio, and omni extensions |
| Core taxonomy | Traditional multimodal models, MLLMs, UMMs, and NMMs with explicit architecture/training qualifications |
| Key distinction | U+G unification for UMMs; architectural and optimization coupling for NMMs, with overlapping membership allowed |
| What makes this repo different | Architecture-first categorization, fusion-aware definitions, and curated links to adjacent awesome lists |
| Intended audience | Researchers, students, and engineers building or surveying multimodal systems |
| Principle | Rule |
|---|---|
| Source quality | Prefer official conference proceedings, OpenReview, ACL Anthology, CVF Open Access, arXiv, and official project pages |
| Classification policy | Category assignment is based on this repository's architecture-first definitions, which may differ from authors' own branding |
| Venue policy | If a peer-reviewed venue is known, we list that venue; otherwise we keep the entry as arXiv |
| Scope discipline | Models, benchmarks, datasets, and analysis papers are tracked separately to avoid mixing artifacts |
| Inclusion bar | We prioritize landmark papers, broadly adopted benchmarks, open implementations, or papers that clarify important taxonomy boundaries |
Classification note: for ambiguous models sitting between
MLLM,UMM, andNMM, this list records the category that best matches the training recipe and architectural coupling, not just the paper title. Architecture-native and training-native qualifications distinguish fusion structure from training history.
In this section: 1.1 Multimodal Model Evolution Stages · 1.2 Scope & Taxonomy · 1.3 Architecture Diagrams
Subtopics: Traditional Multimodal Models · Multimodal Large Language Models (MLLMs) · Unified Multimodal Models (UMMs) · Native Multimodal Models (NMMs)
We use the following precise, architecture-first definitions throughout this list. Understanding these distinctions is critical for correctly classifying modern models.
Pre-2023 mainstream era
Independent per-modality processing followed by simple fusion (early, late, or hybrid). No large-scale language model backbone. Focuses on representation alignment, cross-modal retrieval, and captioning. Examples: CLIP, ALIGN, ViLBERT, BLIP.
Pretrained-backbone multimodal language models
Combine a pretrained visual backbone or visual abstractor (e.g., ViT/CLIP/SigLIP, Q-Former, cross-attention adapter) with a pretrained LLM through a connector. The defining property is inheritance from strong pretrained unimodal backbones rather than joint multimodal pretraining from scratch. These models are primarily text-output understanding/reasoning systems, even when auxiliary generators are attached externally.
Key characteristics:
- ✅ Pretrained visual encoder / abstractor
- ✅ Pretrained LLM backbone
- ✅ Connector layer or cross-attention bridge
- ❌ No end-to-end multimodal pretraining from scratch
- ❌ No native image generation inside the same backbone
Examples: LLaVA, Qwen-VL, InternVL, MiniCPM-V, CogVLM
Single framework for Understanding + Generation (U+G)
A single framework that handles both multimodal understanding and visual generation. UMMs may reuse pretrained components or modular tokenizers; the defining feature is U+G unification, not whether the model is trained from scratch.
Key characteristics:
- ✅ Unified understanding + generation
- ✅ Shared model interface or shared backbone for U+G
⚠️ May use pretrained components⚠️ May use decoupled encoders / modular tokenizers⚠️ If a model is also natively trained from scratch, its architectural details belong primarily in NMMs (§5)
Examples: Show-o, Janus, OpenUni, BAGEL, BLIP3-o
Architectural integration and multimodal optimization
NMMs are described along two dimensions: architectural nativity captures how early and persistently modalities share core computation; optimization nativity captures when multimodal objectives enter training and which core parameters they optimize. Entries that establish only one dimension are qualified as architecture-native or training-native. Training from scratch provides strong optimization evidence, but is not assumed for every entry. UMM and NMM membership can overlap.
Key characteristics:
- ✅ Fusion structure and training history are recorded separately
- ✅ Pretrained components and initialization are annotated per model
- ✅ Architecture-native entries may inherit pretrained backbones
- ✅ Input: text tokens + image patches/tokens
- ✅ Output: text (understanding focus; generation optional)
NMMs are further described by fusion architecture; training-native entries without a confirmed fusion classification are listed separately in §5.5:
Modality tokens or embeddings enter shared core computation before the first shared backbone block. Separate tokenizers, input encoders, modality-specific experts, or output decoders can remain; their presence does not by itself imply late fusion.
- Shared backbone computation over multimodal states
- Discrete tokens, continuous latents, or pixel patches
- Modality interaction from the first shared block
- Initialization and trainable components annotated separately
- Examples: Chameleon, Emu3, Transfusion; Show-o, Show-o2, OneCat, and Llama 4 carry architecture-native qualifications
Each modality is first processed by a dedicated unimodal component (e.g., a vision tower or image encoder), but these components are jointly trained from scratch (not pretrained). Cross-modal interaction occurs at deeper layers.
- Separate unimodal processing stages (trained from scratch)
- Cross-modal interaction at deeper layers
- More modality-specific parameters
- Examples: Models with jointly-trained vision encoders → decoder interaction
Multimodal Models
├── 2. Traditional Multimodal Models
│ ├── 2.1 Multimodel Representations & Alignment
│ │ ├── Multimodal Representations
│ │ ├── Multimodal Fusion
│ │ └── Multimodal Alignment
│ └── 2.2 Multimodal Pretraining
├── 3. Multimodal Large Language Models (MLLMs)
│ ├── 3.1 Foundation MLLMs
│ └── 3.2 Omni MLLMs
├── 4. Unified Multimodal Models (UMMs)
│ ├── 4.1 Taxonomy by Generation Paradigm
│ │ ├── Diffusion-Based UMMs
│ │ ├── Autoregressive (AR) UMMs
│ │ │ ├── Pixel Encoding
│ │ │ ├── Semantic Encoding
│ │ │ ├── Learnable Query Encoding
│ │ │ ├── Hybrid Encoding (Pseduo)
│ │ │ └── Hybrid Encoding (Joint)
│ │ └── Hybrid (AR + Diffusion) UMMs
│ │ ├── Pixel Encoding
│ │ └── Hybrid Encoding
│ └── 4.2 Any-to-Any / Omni UMMs
└── 5. Native Multimodal Models (NMMs)
├── 5.1 Design Analyses & Scaling Laws
├── 5.2 Early Fusion NMMs
├── 5.3 Late Fusion NMMs
├── 5.4 Any-to-Any / Omni NMMs
└── 5.5 Training-Native Models
The NMM diagrams below illustrate from-scratch variants; initialization and input interfaces are specified per model in §5.
┌─────────────────────────────────────────────────────────────────┐
│ TRADITIONAL MULTIMODAL MODEL │
│ │
│ [Image] ──► [CNN/ViT Encoder] ──┐ │
│ ├──► [Fusion] ──► [Output] │
│ [Text] ──► [LSTM/BERT] ──┘ │
└─────────────────────────────────────────────────────────────────┘
┌─────────────────────────────────────────────────────────────────┐
│ MLLM — MODULAR LATE FUSION │
│ │
│ [Image] ──► [Pretrained ViT/CLIP] ──► [Projector/Q-Former] │
│ │ │
│ ▼ │
│ [Text] ──────────────────────────► [Pretrained LLM] ──► [Text]│
└─────────────────────────────────────────────────────────────────┘
┌─────────────────────────────────────────────────────────────────┐
│ UMM — UNIFIED UNDERSTANDING + GENERATION │
│ │
│ [Image/Text Input] ──► [Shared/Modular Tokenizer] │
│ │ │
│ ▼ │
│ [Unified Transformer] │
│ │ │ │
│ ▼ ▼ │
│ [Text Output] [Image Output] │
└─────────────────────────────────────────────────────────────────┘
┌─────────────────────────────────────────────────────────────────┐
│ NMM — EARLY FUSION (Trained from Scratch) │
│ │
│ [Text tokens] ──┐ │
│ └──► [Single Decoder Transformer] ──► [Text] │
│ [Image patches ──► Linear Patchify] ──┘ │
│ (raw pixels, minimal preprocessing) │
│ Multimodal interaction from Layer 1 │
└─────────────────────────────────────────────────────────────────┘
┌─────────────────────────────────────────────────────────────────┐
│ NMM — LATE FUSION (Trained from Scratch) │
│ │
│ [Image] ──► [Jointly-Trained Vision Component] │
│ │ │
│ ▼ (deep layers) │
│ [Text] ──────────► [Cross-Modal Interaction] ──► [Text] │
│ (All components trained jointly from scratch) │
└─────────────────────────────────────────────────────────────────┘
In this section: 2.1 Multimodel Representations & Alignment · 2.2 Multimodal Pretraining
Pre-chat-MLLM and non-native multimodal systems that established the basic vocabulary of alignment, fusion, retrieval, captioning, and multimodal pretraining.
Subtopics: Multimodal Representations · Multimodal Fusion · Multimodal Alignment
| Paper | Venue | Links | Notes | Task |
|---|---|---|---|---|
| Identifiability Results for Multimodal Contrastive Learning | ICLR 2023 | Paper | Theoretical identifiability analysis of contrastive multimodal learning | representation learning |
| Unpaired Vision-Language Pre-training via Cross-Modal CutMix | ICML 2022 | Paper | Introduces CutMix-style augmentation for unpaired VLP | vision-language pretraining |
| Balanced Multimodal Learning via On-the-fly Gradient Modulation | CVPR 2022 | Paper | Balances modality learning via dynamic gradient reweighting | multimodal optimization |
| FLAVA: A Foundational Language And Vision Alignment Model | arXiv 2021 | Paper | Unified architecture for vision-language understanding and generation | foundation multimodal model |
| UniT: Multimodal Multitask Learning with a Unified Transformer | arXiv 2021 | Paper | Single transformer for multiple multimodal tasks | multimodal multitask learning |
| MultiBench: Multiscale Benchmarks for Multimodal Representation Learning | NeurIPS 2021 | Paper | Benchmark suite for multimodal learning evaluation | benchmarking |
| Perceiver: General Perception with Iterative Attention | ICML 2021 | Paper | General-purpose architecture for high-dimensional multimodal inputs | general multimodal architecture |
| Learning Transferable Visual Models From Natural Language Supervision | arXiv 2021 | Paper | Contrastive vision-language pretraining at scale | vision-language contrastive learning |
| VinVL: Revisiting Visual Representations in Vision-Language Models | arXiv 2021 | Paper | Improved visual features for VL tasks | vision-language representation improvement |
| Learning Transferable Visual Models From Natural Language Supervision | arXiv 2020 | Paper | Early large-scale vision-language contrastive learning | vision-language pretraining |
| 12-in-1: Multi-Task Vision and Language Representation Learning | CVPR 2020 | Paper | Unified multi-task learning across 12 VL tasks | multi-task learning |
| Learning Video Representations using Contrastive Bidirectional Transformer | arXiv 2019 | Paper | Contrastive transformer for video representation learning | video contrastive learning |
| OmniNet: A Unified Architecture for Multi-modal Multi-task Learning | arXiv 2019 | Paper | Unified encoder-decoder for multimodal tasks | unified multimodal architecture |
| Learning Representations by Maximizing Mutual Information Across Views | arXiv 2019 | Paper | InfoMax principle for cross-view representation learning | self-supervised learning |
| ViCo: Word Embeddings from Visual Co-occurrences | ICCV 2019 | Paper | Learning word embeddings from visual context | vision-language embeddings |
| Learning Factorized Multimodal Representations | ICLR 2019 | Paper | Factorized latent space for multimodal data | representation disentanglement |
| Deep Fragment Embeddings for Bidirectional Image Sentence Mapping | NeurIPS 2014 | Paper | Fragment-level image-sentence alignment | vision-language alignment |
| DeViSE: A Deep Visual-Semantic Embedding Model | NeurIPS 2013 | Paper | Early deep vision-to-language embedding model | vision-language embedding |
| Multimodal Deep Learning | ICML 2011 | Paper | Foundational multimodal deep learning framework | multimodal deep learning |
| Paper | Venue | Links | Notes | Task |
|---|---|---|---|---|
| Robust Contrastive Learning against Noisy Views | arXiv 2022 | Paper | Robust contrastive learning under noisy multi-view inputs | contrastive learning |
| Attention Bottlenecks for Multimodal Fusion | NeurIPS 2021 | Paper | Introduces bottleneck attention mechanism for efficient multimodal fusion | multimodal fusion |
| VMLoc: Variational Fusion For Learning-Based Multimodal Camera Localization | AAAI 2021 | Paper | Variational multimodal fusion for camera localization tasks | multimodal localization |
| Trusted Multi-View Classification | ICLR 2021 | Paper | Confidence-aware weighting for multi-view classification | multi-view classification |
| Deep-HOSeq: Deep Higher-Order Sequence Fusion for Multimodal Sentiment Analysis | ICDM 2020 | Paper | Higher-order sequence fusion for multimodal sentiment analysis | multimodal sentiment analysis |
| What Makes Training Multi-Modal Classification Networks Hard? | CVPR 2020 | Paper | Analyzes optimization challenges in multimodal classification | theoretical/empirical analysis |
| DeepCU: Integrating Both Common and Unique Latent Information for Multimodal Sentiment Analysis | IJCAI 2019 | Paper | Separates shared and private latent representations for fusion | multimodal sentiment analysis |
| XFlow: Cross-modal Deep Neural Networks for Audiovisual Classification | IEEE TNNLS 2019 | Paper | Cross-modal feature exchange network for audio-visual tasks | audio-visual classification |
| MFAS: Multimodal Fusion Architecture Search | CVPR 2019 | Paper | Neural architecture search for optimal multimodal fusion design | architecture search |
| The Neuro-Symbolic Concept Learner: Interpreting Scenes, Words, and Sentences From Natural Supervision | ICLR 2019 | Paper | Neuro-symbolic model combining perception and reasoning | neuro-symbolic learning |
| Efficient Low-rank Multimodal Fusion with Modality-Specific Factors | ACL 2018 | Paper | Low-rank factorization for efficient multimodal fusion | efficient fusion |
| Memory Fusion Network for Multi-view Sequential Learning | AAAI 2018 | Paper | Memory-based fusion across temporal multimodal sequences | sequential multimodal learning |
| Tensor Fusion Network for Multimodal Sentiment Analysis | EMNLP 2017 | Paper | Tensor-based full interaction modeling across modalities | multimodal sentiment analysis |
| Paper | Venue | Links | Notes | Task |
|---|---|---|---|---|
| LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment | arXiv 2023 | Paper Code | Language-centered alignment across image, video, audio, depth, thermal, and IMU modalities | multimodal alignment |
| CLIP | arXiv 2021 | Paper | 400M+ image-text pairs; dual-encoder (Vision Transformer + Text Transformer); contrastive alignment at embedding level; classic late-fusion foundation | zero-shot image classification, retrieval |
| CoMIR: Contrastive Multimodal Image Representation for Registration | NeurIPS 2020 | Paper | Contrastive learning for multimodal image registration alignment | multimodal alignment |
| Multimodal Transformer for Unaligned Multimodal Language Sequences | ACL 2019 | Paper | Transformer-based alignment for unaligned multimodal sequences | sequence alignment |
| Temporal Cycle-Consistency Learning | CVPR 2019 | Paper | Uses cycle-consistency for temporal cross-modal alignment | temporal alignment |
| Deep Canonical Correlation Analysis | ICML 2013 | Paper | Deep learning extension of CCA for cross-view representation alignment | representation alignment |
| Paper | Venue | Links | Notes | Task |
|---|---|---|---|---|
| Align before Fuse: Vision and Language Representation Learning with Momentum Distillation | NeurIPS 2021 Spotlight | Paper | Momentum distillation for aligning vision-language representations before fusion | vision-language pretraining |
| Less is More: ClipBERT for Video-and-Language Learning via Sparse Sampling | CVPR 2021 | Paper | Sparse frame sampling for efficient video-language pretraining | video-language pretraining |
| UniT: Multimodal Multitask Learning with a Unified Transformer | arXiv 2021 | Paper | Unified transformer for multitask multimodal learning | unified multimodal pretraining |
| Large-Scale Adversarial Training for Vision-and-Language Representation Learning | NeurIPS 2020 | Paper | Adversarial training improves robustness of vision-language representations | robust multimodal pretraining |
| Vokenization: Improving Language Understanding with Contextualized, Visual-Grounded Supervision | EMNLP 2020 | Paper | Grounds language tokens in visual context via voken supervision | vision-grounded language modeling |
| Integrating Multimodal Information in Large Pretrained Transformers | ACL 2020 | Paper | Injects multimodal signals into large pretrained transformer architectures | multimodal transformer pretraining |
| VL-BERT: Pre-training of Generic Visual-Linguistic Representations | arXiv 2019 | Paper | Joint vision-language BERT-style pretraining | vision-language pretraining |
| VisualBERT: A Simple and Performant Baseline for Vision and Language | arXiv 2019 | Paper | Early unified transformer for vision-language understanding | vision-language pretraining |
| ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks | NeurIPS 2019 | Paper | Two-stream transformer for cross-modal vision-language learning | vision-language pretraining |
| Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training | arXiv 2019 | Paper | Cross-modal encoder for universal vision-language representations | vision-language pretraining |
| LXMERT: Learning Cross-Modality Encoder Representations from Transformers | EMNLP 2019 | Paper | Cross-modality transformer encoder for vision-language reasoning | vision-language pretraining |
| VideoBERT: A Joint Model for Video and Language Representation Learning | ICCV 2019 | Paper | Joint discrete token modeling for video and language | video-language pretraining |
In this section: 3.1 Foundation MLLMs · 3.2 Omni MLLMs
Models that connect a pretrained visual encoder / abstractor to a pretrained LLM. Primarily text-output understanding and reasoning systems, defined by inherited pretrained unimodal backbones rather than multimodal pretraining from scratch.
| Paper | Venue | Links | Notes | Task |
|---|---|---|---|---|
| Ling-3.0-flash-VL | Model release 2026 | HF | Open-weight VLM (MIT) extending Ling-3.0-flash with a ViT and a two-layer MLP projector; supports image and video inputs with text outputs. BF16 and quantized checkpoints belong to the same model entry | image/video understanding, document understanding, visual reasoning, GUI agents |
| LLaDA-UI | Model release 2026 | Code HF | SigLIP-initialized native-resolution ViT connects to LLaDA2.0-mini-base through spatial 4-to-1 feature grouping and a two-layer MLP; block-wise diffusion generates text reasoning and GUI actions. Code and weights are available; checkpoint license is unspecified in the model card as of 2026-09-15 | screenshot understanding, GUI grounding, mobile/desktop/web agents |
| Muse Glimmer-30B | Model release 2026 | HF Implementation | Open-weight 30B model (Apache 2.0), distilled from Muse Spark; a dedicated Perception Encoder connects to the language backbone through an MLP adapter and linear projection; text and image inputs, text outputs | visual understanding, visual reasoning, coding, agentic tool use |
| DiffusionGemma | Google DeepMind, 2026 | Model Card HF Implementation | Open-weight Gemma 4-based model (Apache 2.0); normalized visual features enter the language model through a linear projection, while an encoder-decoder architecture uses block-wise discrete diffusion for text generation. Image/video understanding with text output; no native image generation | visual understanding, video understanding, reasoning, efficient text generation |
| Mistral Medium 3.5 | Mistral AI, 2026 | Announcement HF Implementation | Open-weight 128B dense VLM under a Modified MIT license; a from-scratch vision encoder connects through a patch merger and MLP projector to the language backbone. Supports text/image inputs and text outputs; whole-model from-scratch multimodal training is not established | visual understanding, reasoning, coding, long-horizon agents |
| Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model | arXiv 2026 | Paper Project Model | 4B codec-native streaming VLM combining a from-scratch Mage-ViT visual encoder, two-layer MLP projector, and pretrained Qwen3-4B-Instruct-2507 decoder; codec-guided token selection reduces visual tokens by over 75% | image/video understanding, long video, proactive streaming |
| StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design | arXiv 2026 | Paper | 0.9B on-device UI VLM with UI-aware layered visual encoding and a progressive dimensionality projection connector; quantized deployment is validated on Snapdragon 8 Gen5 | UI understanding, OCR, grounding, on-device deployment |
| MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe | arXiv 2025 | Paper Code Model | Efficient 8B MLLM built with Qwen3-8B, SigLIP2-400M and a unified 3D-Resampler for compact image-video encoding | visual understanding, document/OCR, video understanding, edge deployment |
| Moondream 3.1 | Model Release 2026 | Project Model | Efficient sparse MoE VLM with 9B total and 2B active parameters; supports visual reasoning, query, caption, detection, pointing and segmentation | visual understanding, grounding, segmentation, edge deployment |
| LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training | arXiv 2025 | Paper Code HF | 4B / 8B; 85M mid-training data + 22M instruction data; ViT–MLP–LLM architecture with pretrained RICE-ViT, a two-layer MLP projector, and Qwen3 backbone; supports image, multi-image, and video | visual understanding, video |
| SAIL-VL2 Technical Report | arXiv 2025 | Paper | Open-suite 2B/8B vision-language foundation model with SAIL-ViT, progressive multimodal training, SFT-RL thinking fusion, and strong image/video reasoning across 106 datasets | visual understanding, video, reasoning |
| Kwai Keye-VL 2.0 Technical Report | arXiv 2026 | Paper | achieves state-of-the-art performance among models of similar scale, particularly excelling in fine-grained temporal localization | video understanding |
| Penguin-VL: Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders | arXiv 2026 | Paper | LLM-initialized vision encoder (non-CLIP); text-to-vision weight reuse, generative-aligned visual features, optimized for dense perception. | visual understanding |
| Youtu-VL: Unleashing Visual Potential via Unified Vision-Language Supervision | arXiv 2026 | Paper | Tri-modal (V+A+L) unified framework; parameter-efficient tuning, seamless cross-modal reasoning for mobile/IoT deployment. | visual understanding |
| STEP3-VL-10B Technical Report | arXiv 2026 | Paper | 10B-scale foundation multimodal; unified unfrozen pre-training + PaCoRe test-time scaling, frontier-level reasoning with compact footprint. | visual understanding |
| GLM-OCR | arXiv 2026 | Paper | GLM-OCR is an efficient 0.9B-parameter compact multimodal model designed for real-world document understanding. | OCR, structured extraction |
| Kimi K2.5 | arXiv 2026 | Paper | joint text-vision pretraining, Agent Swarm framework; coding, vision, reasoning, agentic tasks; reduces latency by up to 4.5x | visual agentic intelligence, agentic, reasoning |
| Kwai Keye-VL 1.5 Technical Report | arXiv 2025 | Paper | Adaptive Slow-Fast encoding; 8B parameter scale with 128K long-context; SOTA video reasoning & human-preference aligned. | visual understanding |
| olmOCR / olmOCR-2 | arXiv 2025 | Paper | Efficient low-VRAM OCR model based on Qwen2.5-VL fine-tune; excels at preserving semantic structure and markdown output | OCR, structured extraction |
| PaddleOCR-VL | arXiv 2025 | HF / Official | Lightweight (0.9B+) multimodal OCR with 109 languages support; excellent chart-to-HTML/Markdown conversion and high-throughput | OCR, multilingual document |
| DeepSeek-OCR | arXiv 2025 | Paper HF | Lightweight ~3B MoE vision model optimized for high-volume OCR, document digitization, charts and formulas; efficient inference | OCR, document |
| Kimi-VL | arXiv 2025 | Paper HF | Projector + MoE backbone; long video/PDF/GUI, agentic capabilities, chain-of-thought vision reasoning | visual understanding, agentic, video |
| Seed1.5-VL Technical Report | arXiv 2025 | Paper | 20B MoE + 532M ViT; native-resolution vision-language foundation model; efficient asymmetric architecture. | visual understanding |
| Qwen3-VL | arXiv 2025 | Paper HF | Frontier-grade vision/OCR (32+ languages), video analysis, agentic capabilities, strong multimodal reasoning; includes large MoE variants (e.g., 235B) | visual understanding, video, omni |
| SmolVLM | arXiv 2025 | HF | Ultra-lightweight (256M–2.2B) projector-based series; efficient on-device video and image understanding | visual understanding, efficiency |
| LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning | arXiv 2025 | Paper | Diffusion LLM as LLM backbone; vision encoder: SigLIP | visual understanding |
| jina-vlm | arXiv 2025 | Paper HF | SigLIP2 + Qwen backbone with custom projector; optimized for semantic VQA, diagrams, scans and document semantics | visual understanding, VQA, document |
| Phi-4-Multimodal | arXiv 2025 | Paper HF | Small-parameter (LoRA + projectors) multimodal; vision + speech support, efficient on-device deployment | visual understanding, on-device |
| Molmo / PixMo | CVPR 2025 | Paper Code | Strong open-data/open-weight VLM pipeline | visual understanding |
| FastVLM: Efficient Vision Encoding for Vision Language Models | CVPR 2025 | Paper | efficient multimodal visual encoding for on-device deployment | visual understanding, on-device |
| Qwen2.5-VL: Technical Report | arXiv 2025 | Paper HF | Stronger document, grounding, and video capabilities | visual understanding |
| General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model | arXiv 2024 | Paper HF/Code | Specialized end-to-end OCR model with grounding (boxes + points); strong on scientific papers, slides, and mixed visual-text docs | OCR, grounding |
| LLaVA-OneVision: Easy Visual Task Transfer | arXiv 2024 | Paper Code | Single model for image, multi-image, and video transfer | visual understanding |
| MiniCPM-V: A GPT-4V Level MLLM on Your Phone | arXiv 2024 | Paper Code | On-device efficient MLLM | visual understanding |
| NVILA: Efficient Frontier Visual Language Models | CVPR 2025 | Paper | Efficient general purpose multimodal LLM; spatial and temporal "Scale then compress" design; vision encoder: SigLIP | visual understanding |
| GLM-4V | arXiv 2024 | Paper | ViT-based vision encoder (EVA-02-CLIP-L); high-resolution input support (up to 1120x1120) via image tiling, late-fusion architecture, optimized for document and OCR tasks | visual understanding |
| xGen-MM (BLIP-3) | arXiv 2024 | Paper | Open training recipe, datasets, and safety-tuned variants | visual understanding |
| DeepSeek-VL2: Mixture-of-Experts Vision-Language Models | arXiv 2024 | Paper Code | MoE VLM with dynamic tiling and efficient inference | visual understanding |
| Pixtral | arXiv 2024 | Paper HF | 12B open-weight model with strong instruction following, image+text understanding; competitive with larger open VLMs | visual understanding |
| Qwen2-VL | arXiv 2024 | Paper HF | Dynamic resolution; native video | visual understanding |
| Cambrian-1: A Fully Open, Vision-Centric Exploration | NeurIPS 2024 | Paper Code | Spatial Vision Aggregator | visual understanding |
| PaliGemma: A Versatile 3B VLM for Transfer | arXiv 2024 | Paper HF | SigLIP encoder + Gemma backbone; strong transfer model | visual understanding |
| InternLM-XComposer2 | arXiv 2024 | Paper Code | Compositional visual grounding | visual understanding |
| Phi-3-Vision | arXiv 2024 | Paper HF | Small but capable | visual understanding |
| LLaVA-HR: High Resolution MLLMs | CVPR 2024 | Paper | Mixture-of-Resolution Adaptation | visual understanding |
| InternVL2 | Model release 2024 | HF | Instruction-tuned InternVL family release with strong multilingual and OCR capabilities | visual understanding |
| InternVL: Scaling up Vision Foundation Models | CVPR 2024 | Paper Code | Progressively aligned ViT + LLM | visual understanding |
| MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training | arXiv 2024 | Paper | Large-scale proprietary recipe study for multimodal LLM pretraining | visual understanding |
| Ovis: Structural Embedding Alignment for Multimodal Large Language Model | arXiv 2024 | Paper Code | Structural embedding alignment between visual tokens and LLM token space | visual understanding |
| TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document | arXiv 2024 | Paper Code | OCR-free document MLLM emphasizing text-heavy images | document understanding, OCR |
| MoE-LLaVA: Mixture of Experts for Large Vision-Language Models | arXiv 2024 | Paper Code | Sparse MoE extension of LLaVA-style visual instruction tuning | visual understanding |
| MobileVLM: A Fast, Strong and Open Vision Language Assistant for Mobile Devices | arXiv 2023 | Paper Code | Lightweight VLM for mobile deployment | efficient visual assistant |
| Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models | arXiv 2023 | Paper Code | Expands visual vocabulary for dense OCR/document-style perception | document understanding, OCR |
| LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models | arXiv 2023 | Paper Code | Compresses each frame/image into compact context tokens for efficient video MLLMs | video understanding |
| Video-LLaVA: Learning United Visual Representation by Alignment Before Projection | arXiv 2023 | Paper Code | Unified image-video representation before projection into LLM | image/video understanding |
| GLaMM: Pixel Grounding Large Multimodal Model | arXiv 2023 | Paper Code | Pixel-level grounding and phrase-region reasoning | grounded visual understanding |
| Ferret: Refer and Ground Anything Anywhere at Any Granularity | arXiv 2023 | Paper Code | Referring and grounding across points, boxes, and free-form regions | grounded visual dialogue |
| LLaVA-1.5: Improved Baselines with Visual Instruction Tuning | arXiv 2023 | Paper Code | Strong simple baseline with CLIP visual encoder, MLP projector, and instruction tuning | visual instruction tuning |
| ImageBind-LLM: Multi-modality Instruction Tuning | arXiv 2023 | Paper Code | Connects ImageBind-aligned modalities to an LLM for multi-modality instruction following | multi-modal instruction tuning |
| PointLLM: Empowering Large Language Models to Understand Point Clouds | arXiv 2023 | Paper Code | Extends LLM-based multimodal understanding to 3D point clouds | 3D understanding |
| LISA: Reasoning Segmentation via Large Language Model | arXiv 2023 | Paper Code | Couples MLLM reasoning with segmentation mask output | reasoning segmentation |
| GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest | arXiv 2023 | Paper Code | Region-of-interest instruction tuning for fine-grained visual reasoning | region-level understanding |
| 3D-LLM: Injecting the 3D World into Large Language Models | arXiv 2023 | Paper Code | Projects 3D scene features into LLMs for 3D reasoning and dialogue | 3D understanding |
| Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic | arXiv 2023 | Paper Code | Referential dialogue with natural-language coordinates and grounding | grounded visual dialogue |
| LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day | arXiv 2023 | Paper Code | Biomedical adaptation of LLaVA-style visual instruction tuning | biomedical visual assistant |
| DetGPT: Detect What You Need via Reasoning | arXiv 2023 | Paper Code | Uses LLM reasoning to orchestrate detection tools and visual grounding | detection reasoning |
| VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks | arXiv 2023 | Paper Code | Treats vision-centric tasks as open-ended decoding with LLMs | vision-centric decoding |
| MultiModal-GPT: A Vision and Language Model for Dialogue with Humans | arXiv 2023 | Paper Code | Instruction-tuned image-text dialogue model | visual dialogue |
| LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model | arXiv 2023 | Paper Code | Parameter-efficient visual instruction tuning via adapters | visual instruction tuning |
| LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention | arXiv 2023 | Paper Code | Early adapter-based multimodal instruction tuning on LLaMA | visual instruction tuning |
| LLaVA | arXiv 2023 | Paper Code | 7B / 13B+ CLIP vision encoder (frozen/pretrained) + linear projection to LLM (Vicuna/LLaMA); common late-fusion baseline | visual instruction tuning, VQA, image captioning |
| Paper | Venue | Links | Notes | Task |
|---|---|---|---|---|
| M-MiniGPT4: Multilingual VLLM Alignment via Translated Data | arXiv 2026 | Paper | Q-Former based (inherits from MiniGPT-4 / BLIP-2) | vision-language understanding |
| Video Q-Former: Multimodal Large Language Model with Spatio-Temporal Querying Transformer | OpenReview | Paper | Spatio-temporal Q-Former (learnable queries for video spatial-temporal feature extraction) | video understanding |
| HierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video Understanding | arXiv 2025 | Paper | Hierarchical Q-Former (multi-level learnable queries with memory bank for long video) | long video understanding |
| Towards Efficient Visual-Language Alignment of the Q-Former | arXiv 2024 | Paper | PEFT-tuned Q-Former (parameter-efficient fine-tuning on InstructBLIP-style Q-Former) | visual reasoning |
| Matryoshka Query Transformer (MQT) for Large Vision-Language Models | NeurIPS 2024 | Paper | Matryoshka Query Transformer (elastic learnable queries, variable token count) | vision-language understanding |
| Semantically Grounded QFormer for Efficient Vision Language Understanding | arXiv 2023 | Paper | Improved Grounded QFormer (direct latent conditioning, bypass input projection) | vision-language understanding |
| Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding | arXiv 2023 | Paper Code | BLIP-2/Q-Former-style visual and audio query transformers for video dialogue | audio-video understanding |
| InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning | arXiv 2023 | Paper Code | Instruction-aware Q-Former trained over diverse vision-language tasks | visual instruction tuning |
| VideoChat: Chat-Centric Video Understanding | arXiv 2023 | Paper Code | Video-centric MLLM for temporal dialogue and understanding | video understanding |
| mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality | arXiv 2023 | Paper Code | Modular visual abstractor connected to a pretrained LLM | visual instruction tuning |
| MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models | arXiv 2023 | Paper Code | Uses BLIP-2 visual encoder/Q-Former and aligns visual features to Vicuna | visual dialogue |
| BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models | ICML 2023 | Paper Code | Foundational Q-Former architecture bridging frozen vision encoders and frozen LLMs | vision-language pretraining |
| Paper | Venue | Links | Notes | Task |
|---|---|---|---|---|
| CASA: Cross-Attention over Self-Attention | arXiv 2025 | Paper | Efficient cross-attention via self-attention reformulation; competitive with token insertion on image benchmarks, strong for long video | efficient vision-language fusion, video captioning |
| LLaMA 3.2 Vision | arXiv 2024 | Paper HF | Adapter-based vision addition to Llama 3.2; strong OCR, document VQA, 128K context | visual understanding, document |
| Idefics2 | arXiv 2024 | Paper HF | Flamingo-style with Perceiver Resampler + gated cross-attention; improved efficiency on Mistral backbone | open multimodal understanding |
| CogVLM: Visual Expert for Pretrained Language Models | arXiv 2023 | Paper Code | Deep fusion with visual expert modules inside a pretrained LLM | visual understanding |
| Qwen-VL: A Versatile Vision-Language Model | arXiv 2023 | Paper HF | High-res, multi-lang, bounding box | visual understanding |
| Kosmos-2: Grounding Multimodal Large Language Models to the World | arXiv 2023 | Paper Code | Adds grounded phrase-region modeling to multimodal language modeling | grounded visual understanding |
| Kosmos-1: Language Is Not All You Need: Aligning Perception with Language Models | arXiv 2023 | Paper Code | Perception-language model aligning images and language for multimodal reasoning | multimodal reasoning |
| Flamingo: a Visual Language Model for Few-Shot Learning | NeurIPS 2022 | Paper | Perceiver Resampler + gated cross-attention layers for few-shot multimodal prompting | few-shot visual understanding |
| IDEFICS | — | Hugging Face | 80B Flamingo-inspired model; late fusion with vision encoder and LLM | open-source multimodal understanding |
| Paper | Venue | Links | Notes | Task |
|---|---|---|---|---|
| MiniCPM-V 4.6 | Model release 2026 | HF Architecture | Open-weight VLM (Apache 2.0) combining a SigLIP vision encoder, window-attention merger, MLP merger, and Qwen3.5 language backbone; supports 4x and 16x visual downsampling for efficient image/video understanding | visual understanding, document/OCR, video understanding, on-device deployment |
| EXAONE 4.5 Technical Report | arXiv 2026 | Paper Code | Integrates a dedicated visual encoder with the EXAONE 4.0 framework for multimodal pretraining, with strong document understanding and Korean contextual reasoning | visual understanding, document |
| Phoenix-VL 1.5 Medium Technical Report | arXiv 2026 | Paper | 123B multilingual multimodal model continued-pretrained from Mistral Medium 3.1 on localized multimodal and long-context corpora | visual understanding, multilingual |
| DeepSeek-OCR-2 | arXiv 2026 | Paper HF | Optimized for high-volume OCR, document digitization, charts and formulas; efficient inference | OCR, document |
| Ovis2.5 | arXiv 2025 | Paper | Following VET architecture; excellent document understanding and fine-grained quantization | visual understanding, document |
| Ovis2 | arXiv 2025 | HF | Embedding table / projector architecture; excellent document understanding and fine-grained quantization | visual understanding, document |
| MiniMax-01: Scaling Foundation Models with Lightning Attention | arXiv 2025 | Paper | Hybrid Lightning-Softmax Attention; MoE-based (45.9B active) multimodal; 4M long-context with near-zero prefill latency. | visual understanding |
| mPLUG-Owl3 | arXiv 2024 | Paper Code | Long visual sequences | visual understanding |
| Idefics3 | arXiv 2024 | Paper HF | Open-data recipe with strong document understanding | visual understanding |
| NVLM 1.0: Open Frontier-Class Multimodal LLMs | arXiv 2024 | Paper HF | Hybrid multimodal design with strong OCR and reasoning | visual understanding |
| Idefics2 | arXiv 2024 | Paper HF | Fully open; built on Mistral | visual understanding |
| mPLUG-DocOwl 1.5 / 2: Unified Structure Learning for OCR-free Document Understanding | arXiv 2024 | Paper Code | OCR-free document understanding with unified structure learning; excels at long documents and complex layouts | document understanding, OCR |
| Paper | Venue | Links | Notes | Task | Adaptor |
|---|---|---|---|---|---|
| MiMo-V2.5 | Model release 2026 | HF | Open-weight MoE model (MIT) with a dedicated MiMo ViT and a MiMo-Audio-initialized audio encoder; text pretraining is followed by visual/audio MLP projector warmup and multimodal pretraining. Supports image, video, and audio understanding with text outputs | image/video/audio understanding, multimodal reasoning, long-context agents | MLP Projector |
| MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction | arXiv 2026 | Paper Code Model | 9B edge-oriented omni model using Omni-Flow for simultaneous visual/audio perception and speech response; supports proactive full-duplex interaction with less than 12GB memory | vision-language understanding, audio understanding, speech generation, full-duplex live interaction | Hybrid |
| Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence | arXiv 2026 | Paper Code | Efficiency-optimized omni-modal backbone using Hybrid Mamba2-Transformer MoE; supports massive multi-modal contexts (10k+ tokens) for long-video reasoning and agentic GUI navigation on edge devices | omni-modal understanding & reasoning | Hybrid |
| OmniGAIA: Towards Native Omni-Modal AI Agents | arXiv 2026 | Paper Code | Comprehensive benchmark for omni-modal agents with complex multi-hop queries across video, audio, and image; includes OmniAtlas agent with tool-integrated reasoning | omni-modal understanding & reasoning | Native |
| OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention | arXiv 2026 | Paper | Reinforced audio-visual reasoning framework with query intention grounding and modality attention fusion | audio-visual reasoning | Hybrid |
| Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data | arXiv 2025 | Paper Code | MoE-based scaling for omnimodal understanding and generation | omni-modal understanding & generation | MLP Projector |
| Omni-AVSR: Towards Unified Multimodal Speech Recognition with Large Language Models | arXiv 2025 | Paper Code | Unified audio-visual speech recognition using LLMs | audio-visual speech recognition | Hybrid |
| LongCat-Flash-Omni Technical Report | arXiv 2025 | Paper Code | Long-context omni-modal model supporting text and audio generation | long-context omni-modal | Hybrid |
| OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM | arXiv 2025 | Paper Code | Architecture and data enhancements for omni-modal understanding | omni-modal understanding | Hybrid |
| InteractiveOmni: A Unified Omni-modal Model for Audio-Visual Multi-turn Dialogue | arXiv 2025 | Paper Code | Unified model for audio-visual multi-turn dialogue | audio-visual dialogue | Hybrid |
| OneLLM: One Framework to Align All Modalities with Language | CVPR 2024 | Paper | Unified framework aligning eight modalities to language through modality tokenizers and lightweight projectors | all-in-one LLM | Hybrid |
| MoME: Mixture of Matryoshka Experts for Audio-Visual Speech Recognition | NeurIPS 2025 | Paper | Mixture of Matryoshka experts for efficient audio-visual speech recognition | audio-visual speech recognition | Hybrid |
| Qwen3-Omni Technical Report | arXiv 2025 | Paper Code | Omni-modal model with text and audio capabilities (Alibaba/Qwen series) | omni-modal | Native |
| Qwen2.5-Omni Technical Report | arXiv 2025 | Paper Code | Omni-modal technical report with text and audio support (Alibaba/Qwen series) | omni-modal | Hybrid |
| MiniCPM-o 2.6: A GPT-4o Level MLLM for Vision, Speech, and Multimodal Live Streaming on Your Phone | 2025 | Paper Code | On-device GPT-4o level MLLM for vision, speech and multimodal live streaming (OpenBMB) | on-device multimodal live streaming | Hybrid |
| Baichuan-Omni Technical Report | arXiv 2024 | Paper Code | Technical report for Baichuan-Omni (Baichuan Inc.) | omni-modal | Hybrid |
| Baichuan-Omni-1.5 Technical Report | arXiv 2025 | Paper Code | Technical report for Baichuan-Omni 1.5 (Baichuan Inc.) | omni-modal | Hybrid |
| VITA: Towards Open-Source Interactive Omni Multimodal LLM | arXiv 2024 | Paper Code | Open-source interactive omni multimodal LLM | interactive omni multimodal | Hybrid |
| VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction | arXiv 2024 | Paper Code | Real-time vision and speech interaction model | real-time multimodal interaction | Hybrid |
| Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities | NeurIPS 2024 | Paper Code | Open-source GPT-4o style model with vision, speech and duplex capabilities | vision-speech duplex | Hybrid |
| Ola: Pushing the Frontiers of Omni-Modal Language Model with Progressive Modality Alignment | arXiv 2025 | Paper Code | Progressive modality alignment for omni-modal language model | omni-modal alignment | MLP Projector |
| MIO: A Foundation Model on Multimodal Tokens | arXiv 2024 | Paper Code | Foundation model based on multimodal tokens | multimodal tokens | Native |
| EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions | CVPR 2024 | Paper Code | Multimodal model supporting seeing, hearing and emotional speech | emotional multimodal | Hybrid |
| Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model | arXiv 2025 | Paper Code | Simultaneous multimodal interactions with language-vision-speech model | simultaneous multimodal | Hybrid |
| ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and Understanding | arXiv 2025 | Paper Code | Native multimodal LLM focused on 3D generation and understanding | 3D multimodal | Native |
| Pengi: An Audio Language Model for Audio Tasks | arXiv 2023 | Paper Code | Audio-language model using audio representations with a frozen language model for audio captioning, QA, and retrieval-style tasks | audio-language understanding | Hybrid |
| LTU: Listen, Think, and Understand | arXiv 2023 | Paper Code | Audio-oriented multimodal instruction-following model for open-ended audio understanding | audio-language understanding | Hybrid |
In this section: 4.1 Taxonomy by Generation Paradigm · 4.2 Any-to-Any / Omni UMMs
Models that unify multimodal understanding and visual generation within one framework. The defining property is U+G unification, not necessarily training from scratch.
Boundary with NMMs: if a unified model's central contribution is native end-to-end multimodal pretraining from scratch, we document its architectural details primarily in §5 NMMs and keep §4 focused on the unified U+G perspective.
Overview of representative paradigms and architectures of Unified Multimodal Models (UMMs). Source: https://github.com/AIDC-AI/Awesome-Unified-Multimodal-Models
Subtopics: Diffusion-Based UMMs · Autoregressive (AR) UMMs · Hybrid (AR + Diffusion) UMMs
Unified models are categorized according to their core generation mechanism for visual output (while supporting strong multimodal understanding). This taxonomy highlights trade-offs in fidelity, reasoning, efficiency, and training stability.
| Model | Venue | Links | Paradigm | Notes | Task |
|---|---|---|---|---|---|
| LLaDA2.0-Uni | arXiv 2026 | Paper Code | Unified Discrete Diffusion | Unified image generation + understanding base on LLaDA2.0 | visual understanding, visual generation |
| Dual Diffusion | arXiv 2025 | Paper Code | Dual Diffusion | Unified image generation + understanding via bidirectional diffusion | visual understanding, visual generation |
| UniDisc | arXiv 2025 | Paper Code | Unified Discrete Diffusion | Discrete diffusion for multimodal U+G | visual understanding, visual generation |
| MMaDA | arXiv 2025 | Paper Code | Multimodal Large Diffusion LM | Diffusion LM for unified understanding/generation | visual understanding, visual generation |
| FUDOKI | arXiv 2025 | Paper | Discrete Flow-based Unified | Kinetic-optimal velocities for U+G | visual understanding, visual generation |
| Muddit | arXiv 2025 | Paper Code | Unified Discrete Diffusion | Liberating generation beyond T2I | visual understanding, visual generation |
| Lavida-O | arXiv 2025 | Paper Code | Elastic Large Masked Diffusion | Elastic masked diffusion for U+G | visual understanding, visual generation |
| UniModel | arXiv 2025 | Paper | Visual-Only MMDiT Framework | Visual-only unified multimodal U+G | visual understanding, visual generation |
| Model | Venue | Links | Modalities | Notes | Task |
|---|---|---|---|---|---|
| LWM | arXiv 2024 | Paper | video + language | World model on million-length video and language with blockwise ring attention | visual understanding, visual generation |
| Chameleon | arXiv 2024 | Paper Code | image + text | Mixed-modal early-fusion foundation models; token-by-token generation | visual understanding, visual generation |
| ANOLE | arXiv 2024 | Paper Code | image + text | Open autoregressive native LMM for interleaved image-text generation | visual understanding, visual generation |
| MMAR | arXiv 2024 | Paper | image + text | Lossless multi-modal auto-regressive probabilistic modeling | visual understanding, visual generation |
| Orthus | arXiv 2024 | Paper Code | image + text | Autoregressive interleaved image-text generation with modality-specific heads | visual understanding, visual generation |
| SynerGen-VL | arXiv 2024 | Paper | image + text | Synergistic image understanding and generation with vision experts and token folding | visual understanding, visual generation |
| Liquid | arXiv 2024 | Paper Code | image + text | Language models are scalable and unified multi-modal generators | visual understanding, visual generation |
| UGen | arXiv 2025 | Paper | image + text | Unified autoregressive multimodal model with progressive vocabulary learning | visual understanding, visual generation |
| Harmon | arXiv 2025 | Paper Code | image + text | Shared MAR encoder for semantic + fine-grained harmony; SOTA GenEval | visual understanding, visual generation |
| TokLIP | arXiv 2025 | Paper Code | image + text | Marry visual tokens to CLIP for U+G | visual understanding, visual generation |
| Selftok | arXiv 2025 | Paper Code | image + text | Discrete visual tokens for AR / Diffusion / Reasoning | visual understanding, visual generation |
| OneCat | arXiv 2025 | Paper Code | image + text | Pure decoder-only unified U+G | visual understanding, visual generation |
| Uni-X | arXiv 2025 | Paper Code | image + text | Two-end-separated architecture mitigating modality conflict | visual understanding, visual generation |
| Emu3 | arXiv 2024 | Paper Code | image + video + text | Early-fusion native autoregressive model; see NMM / Early Fusion for architectural classification | visual understanding, visual generation |
| Title | Venue | Links | Focus | Task |
|---|---|---|---|---|
| Ming-UniVision: Joint Image Understanding and Generation with a Unified Continuous Tokenizer | arXiv 2025 | Paper Code | Unified continuous tokenizer for joint understanding and generation | visual understanding, visual generation |
| Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents | arXiv 2025 | Paper Code | Bridging MLLMs and diffusion models via patch-level CLIP latents | visual understanding, visual generation |
| Qwen-Image Technical Report | arXiv 2025 | Paper Code | High-quality image generation with strong text rendering | visual generation |
| X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again | arXiv 2025 | Paper Code | RL-enhanced discrete autoregressive unified modeling | visual understanding, visual generation |
| Ovis-U1 Technical Report | arXiv 2025 | Paper Code | 3B unified model for understanding, text-to-image and editing | visual understanding, visual generation |
| UniCode²: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation | arXiv 2025 | Paper | Cascaded large-scale codebooks for unified modeling | visual understanding, visual generation |
| OmniGen2: Exploration to Advanced Multimodal Generation | arXiv 2025 | Paper Code | Versatile open-source unified generation model | visual generation |
| Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations | arXiv 2025 | Paper Code | Text-aligned discrete semantic representations | visual understanding, visual generation |
| UniFork: Exploring Modality Alignment for Unified Multimodal Understanding and Generation | arXiv 2025 | Paper Code | Y-shaped architecture for modality alignment | visual understanding, visual generation |
| UniWorld: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation | arXiv 2025 | Paper Code | High-resolution semantic encoders | visual understanding, visual generation |
| Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation | arXiv 2025 | Paper | Auto-regressive foundation model | visual understanding, visual generation |
| DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies | arXiv 2025 | Paper | Dual visual vocabularies | visual understanding, visual generation |
| UniTok: A Unified Tokenizer for Visual Generation and Understanding | arXiv 2025 | Paper Code | Unified tokenizer | visual understanding, visual generation |
| QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation | arXiv 2025 | Paper Code | Text-aligned visual tokenization | visual understanding, visual generation |
| MetaMorph: Multimodal Understanding and Generation via Instruction Tuning | arXiv 2024 | Paper | Instruction tuning for unified multimodal | visual understanding, visual generation |
| ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance | arXiv 2024 | Paper | Self-enhancing unified see-and-draw | visual understanding, visual generation |
| PUMA: Empowering Unified MLLM with Multi-granular Visual Generation | arXiv 2024 | Paper Code | Multi-granular visual generation | visual understanding, visual generation |
| VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation | ICLR 2024 | Paper Code | Unified foundation model | visual understanding, visual generation |
| Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models | arXiv 2024 | Paper Code | Multi-modality potential mining | visual understanding, visual generation |
| MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer | arXiv 2024 | Paper Code | Interleaved image-text generative modeling | visual understanding, visual generation |
| VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation | arXiv 2023 | Paper | Generative pre-trained transformer | visual understanding, visual generation |
| Generative Multimodal Models are In-Context Learners (Emu2) | CVPR 2024 | Paper | In-context learning generative multimodal model | visual understanding, visual generation |
| DreamLLM: Synergistic Multimodal Comprehension and Creation | ICLR 2023 | Paper | Synergistic multimodal comprehension and creation | visual understanding, visual generation |
| LaVIT: Unified Language-Vision Pretraining in LLM with Dynamic Discrete Visual Tokenization | ICLR 2023 | Paper Code | Dynamic discrete visual tokenization | visual understanding, visual generation |
| Emu: Generative Pretraining in Multimodality | ICLR 2023 | Paper | Generative pretraining in multimodality | visual understanding, visual generation |
| Title | Venue | Links | Focus | Task |
|---|---|---|---|---|
| Skywork UniPic 2.0: Building Kontext Model with Online RL for Unified Multimodal Model | arXiv 2025 | Paper | Kontext model with online RL and MetaQuery connector for unified multimodal framework | visual understanding, visual generation, editing |
| TBAC-UniImage: Unified Understanding and Generation by Ladder-Side Diffusion Tuning | arXiv 2025 | Paper | Ladder-side diffusion tuning integrating MLLM and DiT via layer-wise alignment | visual understanding, visual generation |
| UniLiP: Adapting CLIP for Unified Multimodal Understanding, Generation and Editing | arXiv 2025 | Paper | Adapting CLIP with unified continuous tokenizer for reconstruction, generation and editing | visual understanding, visual generation, editing |
| OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation | arXiv 2025 | Paper Code | Simple baseline with learnable queries and lightweight connector bridging MLLM and diffusion | visual understanding, visual generation |
| BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset | arXiv 2025 | Paper | Fully open unified multimodal models with complete architecture, training recipe and datasets | visual understanding, visual generation |
| Ming-Lite-Uni: Advancements in Unified Architecture for Natural Multimodal Interaction | arXiv 2025 | Paper | Unified visual generator and native multimodal autoregressive model for natural interaction | visual understanding, visual generation |
| Nexus-Gen: A Unified Model for Image Understanding, Generation, and Editing | arXiv 2025 | Paper Code | Prefilled autoregression in shared embedding space unifying understanding, generation and editing | visual understanding, visual generation, editing |
| Transfer between Modalities with MetaQueries | arXiv 2025 | Paper Code | Learnable MetaQueries as efficient interface between autoregressive MLLMs and diffusion models | visual understanding, visual generation |
| SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation | arXiv 2024 | Paper Code | Unified multi-granularity visual semantics for arbitrary-size comprehension and generation | visual understanding, visual generation |
| Making LLaMA SEE and Draw with SEED Tokenizer | ICLR 2023 | Paper Code | SEED tokenizer enabling LLaMA for scalable multimodal autoregression (see and draw) | visual understanding, visual generation |
| Planting a SEED of Vision in Large Language Model | arXiv 2023 | Paper Code | SEED image tokenizer with 1D causal dependency and high-level semantics for LLM vision | visual understanding, visual generation |
| Title | Venue | Links | Focus | Task |
|---|---|---|---|---|
| Skywork UniPic: Unified Autoregressive Modeling for Visual Understanding and Generation | arXiv 2025 | Paper Code | Unified autoregressive modeling with decoupled encoding for image understanding, generation and editing | visual understanding, visual generation |
| MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO | arXiv 2025 | Paper Code | Unified VLM with reasoning generation via Reinforcement Learning (RGPO) | multimodal understanding, reasoning generation |
| UniFluid: Unified Autoregressive Visual Generation and Understanding with Continuous Tokens | arXiv 2025 | Paper | Unified autoregressive framework using continuous visual tokens | visual understanding, visual generation |
| OmniMamba: Efficient and Unified Multimodal Understanding and Generation via State Space Models | arXiv 2025 | Paper Code | Efficient linear-time unified multimodal model based on Mamba (state space models) | multimodal understanding, visual generation |
| Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling | arXiv 2025 | Paper Code | Scaled-up version of Janus with improved training strategy, more data and larger model size | multimodal understanding, visual generation |
| Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation | arXiv 2024 | Paper Code | Decoupling visual encoding to enable unified understanding and generation in an autoregressive framework | multimodal understanding, visual generation |
| Title | Venue | Links | Focus | Task |
|---|---|---|---|---|
| UniAR: Unified Multimodal Autoregressive Modeling with Shared Context-Visual Tokenizer is Key to Unification | ICML 2026 | Paper | Shared context-visual tokenizer bridging understanding and generation; AR visual-token prediction with diffusion-based decoding for high-fidelity generation and editing | visual understanding, visual generation, editing |
| AToken: A Unified Tokenizer for Vision | arXiv 2025 | Paper Code | AToken unified visual tokenizer achieving high-fidelity reconstruction and semantic understanding for images, videos and 3D | visual understanding, visual generation |
| UniWeTok: An Unified Binary Tokenizer with Codebook Size 2128 for Unified Multimodal Large Language Model | arXiv 2026 | Paper | UniWeTok unified binary tokenizer with 2^{128} codebook, pre-post distillation and generative-aware prior for MLLMs | visual understanding, visual generation |
| Towards Scalable Pre-training of Visual Tokenizers for Generation | arXiv 2025 | Paper Code | VTP unified visual tokenizer pre-training framework with joint image-text contrastive, self-supervised and reconstruction losses | visual understanding, visual generation |
| The Prism Hypothesis: Harmonizing Semantic and Pixel Representations via Unified Autoencoding | arXiv 2025 | Paper Code | Prism Hypothesis and unified autoencoding (UAE) harmonizing semantic and pixel representations across modalities | visual understanding, visual generation |
| Show-o2: Improved Native Unified Multimodal Models | arXiv 2025 | Paper Code | Improved native unified multimodal models with autoregressive modeling and flow matching for understanding and generation | multimodal understanding and generation |
| UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding | CVPRW 2025 | Paper Code | Unified visual encoding combining discrete and continuous representations for autoregressive multimodal models | multimodal understanding and generation |
| VARGPT-v1.1: Improve Visual Autoregressive Large Unified Model via Iterative Instruction Tuning and Reinforcement Learning | arXiv 2025 | Paper Code | Enhanced visual autoregressive unified model with iterative instruction tuning and DPO reinforcement learning | visual understanding, generation and editing |
| ILLUME+: Illuminating Unified MLLM with Dual Visual Tokenization and Diffusion Refinement | arXiv 2025 | Paper Code | Dual visual tokenization and diffusion refinement for unified multimodal large language model | multimodal understanding and generation |
| SemHiTok: A Unified Image Tokenizer via Semantic-Guided Hierarchical Codebook for Multimodal Understanding and Generation | arXiv 2025 | Paper | Semantic-guided hierarchical codebook for unified image tokenization supporting understanding and generation | multimodal understanding and generation |
| VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model | arXiv 2025 | Paper Code | Visual autoregressive framework unifying understanding and generation in a single MLLM | visual understanding and generation |
| TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation | CVPR 2025 | Paper Code | Unified image tokenizer with dual-codebook architecture bridging understanding and generation | multimodal understanding and generation |
| MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding | arXiv 2024 | Paper | Semantic discrete encoding for unified vision-language model enabling efficient multimodal understanding and generation | multimodal understanding and generation |
| Paper | Venue | Links | Notes | Task |
|---|---|---|---|---|
| Tuna: Taming Unified Visual Representations for Native Unified Multimodal Models | arXiv 2025 | Paper Code | Native unified multimodal model with cascaded VAE + representation encoder for unified continuous visual representations | multimodal understanding and generation |
| LMFusion: Adapting Pretrained Language Models for Multimodal Generation | arXiv 2024 | Paper | Adapting pretrained LLMs (Llama) for multimodal generation by adding parallel diffusion modules while keeping autoregressive text modeling | multimodal understanding and generation |
| MonoFormer: One Transformer for Both Diffusion and Autoregression | arXiv 2024 | Paper Code | Single shared transformer backbone that handles both autoregressive modeling and diffusion for unified multimodal tasks | visual understanding and generation |
| Show-o: One Single Transformer to Unify Multimodal Understanding and Generation | ICLR 2025 | Paper Code | Unified transformer combining autoregressive and discrete diffusion modeling to flexibly handle mixed-modality inputs/outputs | multimodal understanding and generation |
| Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model | ICLR 2025 | Paper | 7B-scale report | Single multimodal model combining next-token text prediction with image diffusion over mixed discrete/continuous sequences |
| Paper | Venue | Links | Notes | Task |
|---|---|---|---|---|
| Mamoda2.5: Enhancing Unified Multimodal Model with DiT-MoE | arXiv 2026 | Paper Code | Unified AR-Diffusion framework coupling multimodal understanding with a DiT-MoE generation backbone; supports image/video generation and editing, including few-step video editing. Paper and inference code are public; official repository lists model weights as under internal review as of 2026-09-15 | multimodal understanding, image/video generation, image/video editing |
| Vision as Unified Multimodal Generation (SenseNova-Vision) | arXiv 2026 | Paper Model Collection | Fine-tunes a pretrained BAGEL-7B-MoT unified model to express computer-vision tasks through native text, image, or mixed generation without task-specific prediction heads | detection, OCR, keypoints, segmentation, depth, normals, point maps, camera pose |
| Qwen-Image-2.0 Technical Report | arXiv 2026 | Paper | Couples Qwen3-VL as condition encoder with a Multimodal Diffusion Transformer for unified high-fidelity image generation and precise editing | multimodal understanding, image generation, editing |
| S1-Omni-Image: A Unified Model for Scientific Image Understanding, Generation, and Editing | arXiv 2026 | Paper | Builds on S1-VL-32B and injects reasoning hidden states into an image generation module for scientific image understanding, generation and editing | scientific image understanding, generation and editing |
| Omni-Video 2: Scaling MLLM-Conditioned Diffusion for Unified Video Generation and Editing | arXiv 2026 | Paper Code | Connects pretrained MLLMs with video diffusion models through lightweight adapters for unified video generation and editing | video generation and editing |
| InternVL-U: Democratizing Unified Multimodal Models for Understanding, Reasoning, Generation and Editing | arXiv 2026 | Paper Code | 4B unified multimodal model integrating a strong MLLM with an MMDiT-based visual generation head for understanding, reasoning, generation and editing | multimodal understanding, reasoning, generation and editing |
| EMMA: Efficient Multimodal Understanding, Generation, and Editing with a Unified Architecture | arXiv 2025 | Paper Code | Efficient unified architecture with autoencoders, channel-wise concatenation, shared-decoupled networks and MoE for understanding, generation and editing | multimodal understanding, generation and editing |
| HBridge: H-Shape Bridging of Heterogeneous Experts for Unified Multimodal Understanding and Generation | arXiv 2025 | Paper | Asymmetric H-shaped architecture bridging heterogeneous experts with symmetric dense mid-layer connections for unified multimodal modeling | multimodal understanding and generation |
| LightFusion: A Light-weighted, Double Fusion Framework for Unified Multimodal Understanding and Generation | arXiv 2025 | Paper Code | Light-weighted double fusion framework that efficiently integrates pretrained vision-language and diffusion models | multimodal understanding and generation |
| BAGEL: Emerging Properties in Unified Multimodal Pretraining | arXiv 2025 | Paper Code | Open-source foundational decoder-only model pretrained on trillions of interleaved multimodal tokens supporting native understanding and generation | multimodal understanding and generation |
| Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation | arXiv 2025 | Paper | Causal interleaved multi-modal generation framework with deep-fusion, dual vision encoders and multi-modal classifier-free guidance | interleaved multimodal generation |
| JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation | arXiv 2024 | Paper Code | Minimalist framework harmonizing autoregressive LLMs with rectified flow for efficient unified understanding and generation | multimodal understanding and generation |
Models that extend unified understanding + generation beyond text and image to support any-to-any modality conversion (audio, video, speech, etc.). These often build on the paradigms above but emphasize native omni-modal tokenization, long-context handling, and cross-modal generation.
| Model | Paper | Links | Notes | Task |
|---|---|---|---|---|
| Kling-Omni Technical Report | arXiv 2025 | Paper | Unified Diffusion Transformer (DiT) framework with Prompt Enhancer for high-fidelity video generation and reasoning-based editing | Multi-modal visual language (MVL) for unified generation and understanding |
| LongCat-Flash-Omni | arXiv 2025 | Paper Code | Efficient omni model with flash-style acceleration and real-time audio-visual interaction (560B parameters) | any-to-any multimodal generation and understanding |
| Ming-flash-omni 2.0 | Model release 2026-02-11 | HF | Open-weight 100B-total / 6B-active MoE model (MIT) based on Ling-2.0; unifies image, text, video, and audio inputs with image, text, and audio outputs, including image editing and controllable speech/audio/music synthesis | multimodal understanding, image generation/editing, audio generation, streaming video conversation |
| Ming-Flash-Omni | arXiv 2025 | Paper Code | Sparse unified MoE architecture (100B total, 6.1B active) for efficient multimodal perception and generation | any-to-any multimodal perception and generation |
| Qwen3-Omni | arXiv 2025 | Paper Code | Next-gen Qwen omni model with unified modality space, maintaining SOTA across text/image/audio/video | any-to-any multimodal understanding and generation |
| Ming-Omni | arXiv 2025 | Paper Code | Unified multimodal architecture for perception + generation (images, text, audio, video) | any-to-any multimodal tasks |
| M2-Omni | arXiv 2025 | Paper | Extends Omni-MLLM with broader modality support and competitive performance to GPT-4o | any-to-any multimodal modeling |
| Spider | arXiv 2024 | Paper Code | Any-to-many multimodal LLM with flexible output heads for arbitrary modality combinations | multimodal understanding and generation |
| MIO | arXiv 2024 | Paper | Token-level unified multimodal foundation model on discrete multimodal tokens | any-to-any multimodal token modeling |
| X-VILA | arXiv 2024 | Paper | Cross-modality alignment for LLM-based multimodal systems (image/video/audio) | multimodal understanding |
| AnyGPT | arXiv 2024 | Paper Code | Discrete token modeling for unified multimodal generation | any-to-any multimodal generation |
| OmniFlow | CVPR 2025 | Paper | Uses multi-modal rectified flows for any-to-any generation across modalities | any-to-any generation across modalities |
| Video-LaVIT | ICML 2024 | Paper Code | Decoupled visual-motion tokenization for video-language modeling | video understanding and generation |
| Unified-IO 2 | CVPR 2024 | Paper Code | Scales autoregressive multimodal models across modalities | any-to-any multimodal tasks (vision, language, audio, action) |
| NExT-GPT | arXiv 2023 | Paper Code | Any-to-any; encoder+LLM+diffusion decoders | visual understanding, visual generation, omni |
In this section: 5.1 Design Analyses & Scaling Laws · 5.2 Early Fusion NMMs · 5.3 Late Fusion NMMs · 5.4 Any-to-Any / Omni NMMs · 5.5 Training-Native Models
Two dimensions of nativity: architectural integration and multimodal optimization. Architecture-native entries may inherit pretrained backbones; training-native entries are distinguished from claims about fusion depth. From-scratch initialization is stated only where supported. Models may also retain their UMM classification.
What recent arXiv work emphasizes: native multimodality is increasingly defined by end-to-end multimodal pretraining, tokenizer/representation co-design, and scaling strategies that explicitly address the asymmetry between vision and language.
Recent arXiv papers sharpen the definition of NMMs and identify the main bottlenecks in native multimodal pretraining.
| Paper | Venue | Links | Insights |
|---|---|---|---|
| Scaling Native Multimodal Pre-Training From Scratch | arXiv 2026 | Paper | Derives compute and allocation laws for transformer-based VLMs trained from scratch, showing distinct language and multimodal scaling behavior, data-mixture-sensitive multimodal allocation, and positive cross-modal transfer |
| Toward Native Multimodal Modeling: A Roadmap | arXiv 2026 | Paper | The end-to-end pipeline from industrial perspectives from architectural coordination, massive data curation, to full-stack training recipes, inference & deployment, and the comprehensive evaluation for truly native modeling |
| Beyond Language Modeling: An Exploration of Multimodal Pretraining | arXiv 2026 | Paper | Highlights representation autoencoders, vision-language data synergy, and MoE for native pretraining |
| NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints | arXiv 2025 | Paper Code | End-to-end native MLLM scaling shows positive correlation between visual encoder and LLM size under data constraints; optimal meta-architecture balances cost and performance |
| Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training | arXiv 2025 | Paper | reveals that LLMs develop latent visual priors during text-only pre-training, where reasoning-centric data (code and math) builds transferable visual reasoning skills while broad corpora foster perception, enabling models to 'see' before ever processing an image. |
| Scaling Laws for Native Multimodal Models | arXiv 2025 | Paper | Early-fusion NMMs match or outperform late-fusion at low compute; early-fusion needs fewer params; MoE with modality-agnostic routing boosts sparse NMM scaling |
| The Narrow Gate: Localized Image-Text Communication in Native Multimodal Models | arXiv 2024 | Paper | Native models often funnel image-to-text communication through a single post-image token |
Modality tokens or embeddings enter shared backbone computation before its first block. Separate input encoders, tokenizers, experts, and decoders are compatible with early fusion. Architecture-native entries are explicitly qualified where pretrained initialization prevents an equally strong from-scratch claim.
Recent scaling-law evidence suggests early-fusion NMMs are often stronger at lower parameter counts and simpler to deploy when paired with sufficiently strong visual representations.
| Model | Paper | Links | Training Scale | Notes | Task |
|---|---|---|---|---|---|
| DeepSeek-V4.1-Flash | Model release 2026-09-10 | Announcement HF | 552B backbone; 45T multimodal pretraining tokens; 1M context | Open-weight model (MIT) trained from scratch on multimodal data; DeepSeek-ViT and a two-layer MLP produce image embeddings jointly processed with text from the start of language-model pretraining. Shared computation uses a Causal Encoder-Decoder architecture | image/text understanding, visual reasoning, coding, long-context agents |
| Qwen3.8-Flash-Next | Model release 2026-08-26 | Code HF Implementation | 125B backbone / 6B active; 51B n-gram embeddings + 4B MTP | Open-weight experimental architecture preview under Qwen Community License 1.0; architecture-native early fusion inserts image/video embeddings into the shared language-backbone input. Combines Gated DeltaNet, Qwen Sparse Attention, and gated residuals; all-components-from-scratch initialization is unverified | image/video understanding, visual reasoning, coding, long-horizon agents |
| Qwen3.8-27B | Model release 2026-08-14 | Code HF | 27B dense; 262K native context, extensible to 1M | Open-weight architecture-native VLM built on Qwen3.5's early-fusion vision-language architecture; image/video embeddings and text share the language backbone. Training stage is pretraining plus post-training; all-components-from-scratch initialization is unverified | image/video understanding, document understanding, visual reasoning, agentic work |
| Inkling-Small | Model release 2026-07-30 | Announcement Model Card HF | 276B total / 12B active | Open-weight MoE model with a hierarchical image-patch encoder and discrete audio encoding; text, image, and audio representations enter a shared hidden space and are jointly processed by one decoder. Outputs text; fusion classification does not assert all-components-from-scratch initialization | text/image/audio understanding, visual and audio reasoning, coding, agentic tool use |
| Inkling | Thinking Machines Lab, 2026 | Announcement Model Card HF | 975B total / 41B active; 45T multimodal pretraining tokens; 1M context | Open-weight MoE model trained from scratch with encoder-free vision and audio inputs; images use 40×40 patches with a lightweight four-layer hMLP, while audio uses dMel spectrograms, jointly processed with text tokens | text/image/audio understanding, visual and audio reasoning, speech transcription, coding, agentic tool use |
| Gemma 4 12B Unified | Model release 2026-06-03 | Announcement HF | 12B | Open-weight encoder-free architecture (Apache 2.0): lightweight linear layers project raw image patches and audio waveforms into the shared decoder's embedding space. Listed separately from encoder-based Gemma 4 variants; "Unified" refers to input architecture, with text-only output | image/video/audio understanding, multimodal reasoning, on-device agents |
| HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers | arXiv 2026 | Paper | 7B dense model | Native unified multimodal model with a holistic image-video ViT tokenizer, covering image/video understanding, image/video generation, and image editing | image/video understanding, image/video generation, editing |
| NEO-OV | arXiv 2026 | Paper | — | Native one-vision scaling from pixels to words, extending the NEO line toward stronger native visual primitives | vision-language understanding |
| SenseNova-U1.5-8B-MoT | Model release 2026-08-20 | Paper Code HF | — | NEO-unify-based encoder-free MoT model with improved patchify layers, native 4K generation, text rendering, and image editing; grouped with U1 by architecture, with from-scratch initialization unverified | visual understanding, image generation, editing, interleaved generation |
| SenseNova-U1 | arXiv 2026 | Paper Code | — | NEO-Unify-based native unified multimodal model handling understanding, generation, and reasoning in one model | Unified Understanding & Generation |
| HiDream-O1-Image | arXiv 2026 | Paper | — | Natively unified image generative foundation model with pixel-level unified transformer | image generation, unified modeling |
| Tuna-2 | arXiv 2026 | Paper | — | Native unified multimodal model that discards traditional vision encoders in favor of direct pixel embeddings for end-to-end understanding and generation | Unified Understanding & Generation |
| NEO | arXiv 2025 | Paper | — | Native vision-language primitives at scale; paired with reusable components for cost-effective native VLM development | vision-language understanding |
| NEO-Unify | Blog 2025 | Blog | — | Native unified extension of NEO for understanding, generation, and reasoning | Unified Understanding & Generation |
| Emu3.5 | Nature 2026 | Paper Code | Large-scale (trillion+ tokens) | Native world model; next-state prediction on interleaved video/text; Discrete Diffusion Adaptation for efficiency | interleaved generation, world modeling, any-to-image |
| LongCat-Next | arXiv 2026 | Paper | — | Early fusion: DiNA represents text, vision, and audio as discrete tokens processed by a shared modality-agnostic autoregressive backbone; Omni capabilities | text/image/audio understanding and generation |
| Llama4 | arXiv 2026 | Paper Blog | Scout/Maverick: 17B active / ~109B–400B total; Behemoth: ~2T total | Architecture-native early fusion of text and vision tokens in a shared MoE backbone; separately trained vision encoder, so early fusion does not imply all components were trained from scratch | vision-language understanding |
| Emu3 | arXiv 2024 | Paper Code | 8B | Early-fusion autoregressive backbone jointly models text, image, and video tokens; also indexed under UMM / AR / Pixel Encoding | visual understanding, visual generation |
| Chameleon | arXiv 2024 | Paper Code | — | Early-fusion backbone trained from scratch on interleaved image-text tokens; separate visual tokenizer; also listed under UMM / AR / Pixel Encoding | visual understanding, visual generation |
| Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model | ICLR 2025 | Paper | 7B | Early-fusion shared Transformer jointly models text tokens and continuous image latents with next-token prediction and diffusion; also listed under Hybrid UMMs | visual understanding, visual generation |
| Show-o: One Single Transformer to Unify Multimodal Understanding and Generation | ICLR 2025 | Paper Code | — | Architecture-native early fusion with a pretrained language-model initialization; shared Transformer combines autoregressive text modeling and masked image-token diffusion; UMM entry | multimodal understanding and generation |
| Show-o2: Improved Native Unified Multimodal Models | arXiv 2025 | Paper Code | — | Architecture-native early fusion with autoregressive text modeling and visual flow matching; uses a pretrained language backbone, not an all-components-from-scratch model; UMM entry | multimodal understanding and generation |
| OneCat | arXiv 2025 | Paper Code | — | Architecture-native early fusion with a Qwen2.5-initialized decoder and modality-specific experts; unified autoregressive understanding and generation; UMM entry | visual understanding, visual generation, editing |
Models where separate unimodal components are jointly trained from scratch (not pretrained), with cross-modal interaction occurring at deeper layers. Distinct from MLLMs where vision encoders are pretrained.
| Model | Paper | Links | Training Scale | Notes | Task |
|---|---|---|---|---|---|
| Lance | arXiv 2026 | Paper Code | 3B (MoE) | Native multimodal MoE | Unified multimodal understanding and multimodal generation |
| Kimi K3: Open Frontier Intelligence | arXiv 2026 | Paper HF | 2.8T total / 104B active; 1M context | Joint language-vision training from the outset; MoonViT-V2 is trained from scratch and connected to the language backbone through pixel-shuffle downsampling and an MLP projector | image/video understanding, visual reasoning, long-horizon agents |
| Kimi K2.6 | Moonshot AI 2026 | Blog | MoE Architecture: 32B active / 1T total parameters; supports 256K context | Native multimodal MoE with MLA (Multi-head Latent Attention) and MoonViT encoder | Unified multimodal understanding, long-horizon coding, and agent swarms |
| GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents | arXiv 2026 | Paper | — | Native multimodal foundation model featuring CogViT vision encoder, Multimodal Multi-Token Prediction (MMTP), and joint RL for agentic GUI/Design2Code tasks | multimodal agentic reasoning |
| LLaVA-OneVision-2 | arXiv 2026 | Code HF | 8B | Native multimodal training with OneVision-Encoder and Qwen3 text backbone; codec-aligned visual encoding for image, long-video, spatial, document/OCR/chart understanding | visual understanding, video, spatial reasoning |
| Qwen3.5 | - | Blog | — | Discrete Native Any-resolution Visual Transformer | vision-language understanding |
| Gemma4 | - | Blog | — | A pre-trained ViT encoder with a visual expert that uses cross-attention for deep but late-style fusion to the LLM, preserving its capabilities. | vision-language understanding |
The latest arXiv-native multimodal papers increasingly blur the boundaries between omni understanding, any-to-any generation, world modeling, and RL-enhanced post-training.
| Model | Paper | Links | Training Scale | Notes | Task |
|---|---|---|---|---|---|
| HyperCLOVA X 8B Omni | arXiv 2026 | Paper HF | 8B | Any-to-any omnimodal model unifying text, audio, and vision through a shared next-token prediction interface over interleaved multimodal sequences | text/audio/vision understanding and generation |
| MiniMax-M3 | MiniMax 2026 | Blog Code | ~428B total / ~23B active; 1M context | Native multimodal model with MiniMax Sparse Attention (MSA), trained for text, image, and video understanding, long-context coding, agentic workflows, and computer-use tasks | native multimodal understanding, video, coding, agentic computer use |
| Tri-Modal Masked Diffusion Models | arXiv 2026 (Omni / Any-to-Any) | Paper | 3B; 6.4T tokens | Studies a from-scratch tri-modal masked diffusion model spanning text, image-text, and audio-text data, with scaling, modality mixing, noise-schedule, batch-size, and inference analyses | text generation, text-to-image, text-to-speech |
| Qwen3.5-Omni | Qwen Blog 2026 | Blog | — | Discrete native any-resolution visual transformer with omni-modal extension | vision-language understanding, omni |
| ERNIE 5.0 Technical Report | arXiv 2026 (Late fusion) | Paper | — | Natively autoregressive foundation model designed for unified multimodal understanding and generation across text, image, video, and audio | vision-language understanding, omni |
These entries record a native multimodal training route separately from fusion topology. This qualification does not claim that every component is trained from scratch; no Early, Mid, or Late Fusion label is assigned without architectural evidence.
| Model | Paper | Links | Training Scale | Notes | Task |
|---|---|---|---|---|---|
| InternVL3 | arXiv 2025 | Paper | — | Training-native route: native multimodal pretraining within a ViT–MLP–LLM architecture; pretrained components are retained. This training qualification does not establish strict late or mid fusion; fusion depth is left unspecified here. | vision-language understanding |
| InternVL3.5 | arXiv 2025 | Paper | — | Training-native route: native multimodal pretraining within a ViT–MLP–LLM architecture; pretrained components are retained. This training qualification does not establish strict late or mid fusion; fusion depth is left unspecified here. | vision-language understanding |
| Model | Venue | Links | Notes | Task |
|---|---|---|---|---|
| GPT-6 Astra | OpenAI, 2026 | Announcement System Card | September 2026 proprietary model with text/image input and text output, supporting visual reasoning, coding, computer use, and long-horizon agentic work | visual understanding, multimodal reasoning, coding, computer use, agents |
| Claude Fable 5.1 / Mythos 5.1 | Anthropic, 2026 | Announcement Models | September 2026 release; Fable 5.1 and Mythos 5.1 are the same underlying model with different safeguards. Fable is generally available, while Mythos is restricted to trusted access programs; grouped as one entry | visual understanding, reasoning, coding, computer use, long-horizon agents |
| Muse Spark 1.3 | Meta, 2026 | Announcement Models | September 2026 proprietary Muse Spark update for multimodal perception, coding, and agentic work, available through Muse Code and Meta Model API. Open-weight release remains a future plan as of 2026-09-15 | visual understanding, multimodal reasoning, coding, agents |
| Gemini 3.8 Flash | Google DeepMind, 2026 | Model Card | September 2026 Gemini Flash model supporting text, image, audio, and video inputs with text output; up to 1M input context, with configurable reasoning effort | multimodal understanding, video/audio understanding, reasoning, coding, agents |
| Grok 4.6 | xAI, 2026 | Announcement Model Docs | August 2026 proprietary model with text/image input and text output; extends the Grok family with visual work, coding, and long-running agent workflows | visual understanding, multimodal reasoning, coding, agents |
| Claude 5 Family (Opus 5 / Sonnet 5) | Anthropic, 2026 | Opus 5 Sonnet 5 Models | Opus 5 (July 2026) and Sonnet 5 (June 2026) support text/image input and text output, with vision, reasoning, tool use, and agentic coding; retained separately from the Fable/Mythos family | visual understanding, reasoning, coding, computer use, agents |
| Seed2.1 Pro / Turbo | ByteDance Seed, 2026 | Announcement Models | June 2026 proprietary Seed2.1 family; Pro and Turbo provide different model sizes for visual/video understanding, reasoning, software engineering, and cross-tool agent workflows | visual/video understanding, multimodal reasoning, coding, agents |
| GPT-5.6 Sol / Terra / Luna | OpenAI, 2026 | Announcement | July 2026 general-availability release of the GPT-5.6 family: Sol is the flagship model, Terra balances capability and cost, and Luna prioritizes speed and affordability. | Multimodal Reasoning + Computer Use + Agentic Work |
| GPT-5.4 / GPT-5.5 | OpenAI Blog | GPT-5.4 GPT-5.5 | 2026 GPT-5-series updates with improved reasoning, multimodal capability, and deployment efficiency. | Omni-Modal + Professional/Agentic Workflows |
| Claude 4.6 Family (Opus 4.6 / Sonnet 4.6) | Anthropic Blog | Opus 4.6 Sonnet 4.6 | Claude 4.6 proprietary model family with vision, coding, tool use, and computer-use workflows. | Multimodal + Agentic/Coding/Computer-Use |
| Gemini 3.1 Pro / Gemini 3.5 | Google Blog | Gemini 3.1 Pro Gemini 3.5 | 2026 Gemini 3-series updates with stronger multimodal reasoning, long-context, and action-oriented capabilities. | Frontier Multimodal (text/image/audio/video + reasoning) |
| Model | Venue | Links | Notes | Task |
|---|---|---|---|---|
| Gemini 3 / Gemini 3 Pro | Google Blog | Gemini 3 | Native multimodal Gemini 3 generation with advanced reasoning, coding, and long-context capabilities. | Frontier Multimodal + Reasoning |
| GPT-5 / GPT-4.5 | OpenAI Blog | GPT-5 GPT-4.5 | Proprietary OpenAI GPT-series updates with multimodal reasoning and tool-use support. | Omni-Modal + Advanced Reasoning/Agentic |
| Claude 4 Family (Opus 4 / Sonnet 4) | Anthropic Blog | Claude 4 | Claude 4 model family for coding, agentic tasks, vision, and extended tool use. | Vision + Advanced Reasoning/Agentic Workflows |
| Grok 3 / Grok 4 / Grok 4.1 | xAI Announcement | Grok 3 Grok 4 Grok 4.1 | Proprietary xAI models with multimodal input, real-time platform integration, and reasoning-focused releases. | Multimodal Reasoning + Real-Time Integration |
| Gemini 2.0 / 2.5 (Pro / Flash) | Google Blog | Gemini 2.0 Gemini 2.5 Pro Gemini 2.5 Updates | Native multimodal Gemini 2-series models with long context, video/audio/image understanding, and agentic features. | Advanced Native Multimodal + Agentic |
| Mistral Medium 3 | Mistral AI | Announcement | Proprietary Mistral model offering text and vision capabilities through API/platform deployments. | General Multimodal Tasks |
| Model | Venue | Links | Notes | Task |
|---|---|---|---|---|
| Gemini 1.5 (Pro / Flash) | Google DeepMind Blog | Gemini 1.5 Announcement | Released February 2024. Massive context (>1M tokens), strong long-context multimodal (video, audio, images). Proprietary. | Long-Context Multimodal (video/audio/image/text) |
| Claude 3 Family (Opus / Sonnet / Haiku) | Anthropic Blog | Claude 3 Family | Released March 2024. Strong native vision for images, charts, diagrams, and documents. Proprietary API + Claude.ai. | Vision-Language + Reasoning |
| Claude 3.5 Sonnet | Anthropic Blog | Announcement | Stronger proprietary Claude vision/reasoning model used for images, documents, charts, and coding workflows. | Vision-Language + Reasoning |
| GPT-4o / GPT-4o mini | OpenAI Blog | GPT-4o GPT-4o mini | GPT-4o introduced real-time omni interactions; GPT-4o mini provided a lower-cost multimodal model. | Real-Time Omni-Modal (text/vision/audio) |
| Grok-1.5V / Grok-2 Vision | xAI Announcement | Grok-1.5V Grok-2 | Vision-capable Grok models for image understanding, diagram reasoning, and X platform integration. | Vision-Language |
| Amazon Nova (Pro / Lite / Canvas / Reel) | Amazon Announcement | AWS Blog | Amazon Bedrock model family covering multimodal understanding and image/video generation variants. | Multimodal Understanding + Generation |
| Pixtral Large | Mistral AI | Announcement | Proprietary frontier multimodal Mistral model for image-text understanding through Le Chat/API deployments. | Vision-Language |
| Model | Venue | Links | Notes | Task |
|---|---|---|---|---|
| GPT-4V (Vision) | OpenAI Announcement | GPT-4V System Card | Released September 2023. First widely available multimodal GPT-4 variant. Image + text input, text output. API/ChatGPT access only. | Vision-Language (image understanding, VQA, OCR, document analysis, captioning) |
| Gemini 1.0 (Ultra / Pro / Nano) | Google DeepMind Blog | Gemini Announcement | Released December 2023. Native multimodal from training (text + image + audio + video). Proprietary API + Gemini chatbot. | Native Multimodal Understanding (text/image/audio/video) |
In this section: 7.1 Related Awesome Lists · 7.2 Slides & Survey Papers · 7.3 Code Repositories & Tools
| Repository | Focus | Author |
|---|---|---|
| awesome-multimodal-ml | General multimodal ML | pliang279 |
| Awesome-Multimodal-Large-Language-Models | MLLMs + evaluation | BradyFU |
| Awesome-Multimodal-Research | Broad multimodal research | Eurus-Holmes |
| Awesome-Unified-Multimodal-Models | UMMs | ShowLab |
| Awesome-Multimodal-Large-Language-Models | MLLMs | yfzhang114 |
| awesome-foundation-and-multimodal-models | Foundation + multimodal | SkalskiP |
| Awesome-Multimodality | General multimodality | Yutong-Zhou-cv |
| Awesome-Unified-Multimodal | Unified models | Purshow |
| Awesome-Unified-Multimodal | Unified models | AIDC-AI |
| Type | Resource | Notes |
|---|---|---|
| Slides | Native LMM Slides | Ziwei Liu (NTU); concise framing for native multimodal models |
| Survey | A Survey on Multimodal Large Language Models | Broad survey of MLLM architectures, data, and evaluation |
| Report | The Dawn of LMMs: Preliminary Explorations with GPT-4V | Early capability analysis around GPT-4V |
| Survey | Multimodal Foundation Models: From Specialists to General-Purpose Assistants | Broader foundation-model view across multimodal systems |
| Tool | Description | Link |
|---|---|---|
| TorchUMM | Unified evaluation, analysis and post-training toolkit for heterogeneous unified multimodal model architectures, tasks and datasets | Paper Code |
| LMMs-Eval | Unified evaluation harness for multimodal models | Code |
| ImageBench | Live text-to-image benchmark ranking 40+ models on 192 prompts across 6 categories using VLM judges; every generated image is published for inspection | Site Methodology |
| LAVIS | Library for Language-Vision Intelligence (Salesforce) | Code |
| OpenFlamingo | Open reproduction of DeepMind Flamingo | Code |
| xtuner | Efficient fine-tuning for multimodal LLMs | Code |
| LLaMA-Factory | Multimodal instruction tuning framework | Code |
| MMEngine | Foundation for perception research (OpenMMLab) | Code |
| DeepSpeed-VisualChat | Scalable multimodal chat training | Code |
In this section: Validation Rules · Entry Format
We welcome contributions! Please follow these guidelines:
For NMM submissions:
- Record architectural integration and multimodal training history separately
- Identify pretrained backbones, encoders, tokenizers, and the trainable parameters at each stage
- Qualify architecture-native or training-native entries; claim from-scratch training only with supporting evidence
- Assign fusion depth only when supported by the architecture; otherwise use the Training-Native group with fusion unspecified
For UMM submissions:
- Confirm the model handles both image understanding AND image generation
- Note whether pretrained components are used (annotate accordingly)
For MLLM submissions:
- Note which vision encoder is used (must be a pretrained encoder)
- Note which LLM backbone is used (must be a pretrained LLM)
| **Model Name** | [Paper](arxiv_link) [Code](github_link) [HF](huggingface_link) BADGES | Scale | Key contribution / notes |Submit a PR with:
- The paper/model entry in the correct section
- A one-line justification for the chosen category
- Links to paper, code, and/or weights
If this list is useful in your research, please consider citing:
@misc{awesome-multimodal-modeling-2026,
title = {Awesome Multimodal Modeling: From Traditional to Native & Unified},
author = OpenEnvision-Lab,
year = {2026},
url = {https://github.com/OpenEnvision-Lab/Awesome-Multimodal-Model-Traditional-Advanced},
note = {GitHub repository}
}This list is released under the CC0 1.0 Universal license.
The companion research library is designed for openenvision.github.io/Awesome-Multimodal-Modeling/, under the OpenEnvision organization site.
- 📄 Survey paper — Coming soon. A survey of multimodal foundation models is in preparation.
- 💻 GitHub repository. Explore the list, suggest entries, and contribute.
The website derives its catalog and category memberships from this README. It supports search, nested category filters, year and resource filters, sorting, pagination, entry details, shareable filter URLs, and a browser-local reading list. Counts represent entries, including cross-listed models and analysis papers, rather than unique models.
index.html lives in the repository root, beside README.md. It is a complete static page,
with its styles, scripts, icons, diagrams, and catalog in assets/site/.
Awesome-Multimodal-Modeling/
├── index.html # Website entry point, ready to serve
├── README.md # Source catalog and project documentation
├── .nojekyll # Serve static files directly on GitHub Pages
├── robots.txt
├── sitemap.xml
├── assets/
│ ├── site/ # Website styles, scripts, icons, diagrams, and catalog.json
│ └── ... # Existing README illustrations
├── scripts/ # Catalog synchronization and local preview tools
└── .github/workflows/pages.yml
Publish index.html together with assets/; keep their relative paths intact.
No server, package installation, or API key is required to host the prepared files.
- Commit the root
index.html,README.md,assets/,.nojekyll,robots.txt, andsitemap.xmltomain. - Open Settings → Pages → Build and deployment.
- Choose Deploy from a branch, select main and /(root), then save.
The site will be available at openenvision.github.io/Awesome-Multimodal-Modeling/. The custom workflow detects branch publishing and skips its own deployment.
After editing the model list, run:
npm test
npm run buildNode.js 22 or later is sufficient. The build refreshes the generated catalog blocks inside the root
index.html, writes assets/site/catalog.json, and exports a clean copy to dist/.
Commit the refreshed root files when publishing from a branch. Page content outside the marked
catalog blocks can be edited directly in the root index.html.
For automatic catalog synchronization on every push, choose GitHub Actions as the Pages source
and keep the included workflow. It checks pull requests, rebuilds the catalog from README changes,
and deploys successful main builds. A first deployment can also be started with Run workflow.
The generated dist/ directory is only used as the Actions deployment artifact and is excluded
from version control.
See GitHub's publishing-source documentation.
npm run previewOpen http://127.0.0.1:4173/Awesome-Multimodal-Modeling/.
This serves the prepared files directly from the repository root.
Use npm run dev for an in-memory preview that watches README.md, index.html, and assets/;
reload the browser after edits.
The project uses relative asset URLs. It needs no custom domain, CNAME file, or change to the OpenEnvision organization's main website.
The brand mark comes from the OpenEnvision site.
The GitHub icon comes from GitHub Octicons;
its MIT license is included in assets/site/OCTICONS-LICENSE.txt.
Model membership, release annotations, and availability qualifications remain governed by the README.
