Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Awesome Unified Multimodal Models
| Date | Stars |
|---|---|
| 2026-07-24 | 1305 |
| 2026-07-25 | 1305 |
| 2026-07-28 | 1306 |
| 2026-07-30 | 1306 |
| 2026-08-06 | 1306 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
[//]: # "# Awesome Unified Multimodal Models" [//]: # [//]: # "## Survey " <div align="center"> <h1 style="font-size: 3em; margin-bottom: 0.2em;">Awesome Unified Multimodal Models</h1> </div> <div align="center" style="font-size: 2em;"> <a href="https://arxiv.org/abs/2505.02567" target="_blank">📚Survey</a> • 🤗 <a href="https://huggingface.co/papers/2505.02567" target="_blank">HF Repo</a> </div> ##  *Figure 1: Timeline of Publicly Available and Unavailable Unified Multimodal Models. The models are categorized by their release years, from 2023 to 2025. Models underlined in the diagram represent any-to-any multimodal models, capable of handling inputs or outputs beyond text and image, such as audio, video, and speech. The timeline highlights the rapid growth in this field.* ## 🔥 We are hiring! We are looking for both interns and full-time researchers to join our team, focusing on multimodal understanding, generation, reasoning, AI agents, and unified multimodal models. If you are interested in exploring these exciting areas, please reach out to us at [email protected]. ## 👉 What is This Repo for? This repository provides a comprehensive collection of resources related to unified multimodal models, featuring: - A survey of advances, challenges, and timelines for unified models - Categorized lists of diffusion-based, autoregressive (MLLM), and hybrid architectures for unified image–text understanding and generation - Benchmarks for evaluating multimodal comprehension, image generation, and interleaved image–text tasks - Representative datasets covering multimodal understanding, text-to-image synthesis, image editing, and interleaved interactions Designed to help researchers and practitioners explore, compare, and build state-of-the-art unified multimodal systems. ## Awesome Papers & Datasets - [Text-and-Image Unified Models](#text-and-image-unified-models) - [Diffusion](#diffusion) - [MLLM AR](#mllm-ar) - [MLLM AR-Diffusion](#mllm-ar-diffusion) - [Any-to-Any Multimodal models](#any-to-any-multimodal-models) - [Benchmark for Evaluation](#benchmark-for-evaluation) - [Benchmarks on Understanding Tasks](#benchmarks-on-understanding-tasks) - [Benchmarks on Image Generation Tasks](#benchmarks-on-image-generation-tasks) - [Benchmarks on Interleaved Tasks](#benchmarks-on-interleaved-tasks) - [Dataset](#dataset) - [Multimodal Understanding](#multimodal-understanding) - [Text-to-Image](#text-to-image) - [Image Editing](#image-editing) - [Interleaved Image-Text](#interleaved-image-text) - [Other Text+Image-to-Image](#other-text-image-to-image) - [Applications and Opportunities](#applications-and-opportunities) ### Text-and-Image Unified Models  *Figure 2: Classification of Unified Multimodal Understanding and Generation Models. The models are divided into three main categories based on their backbone architecture: Diffusion, MLLM (AR), and MLLM (AR + Diffusion). Each category is further subdivided according to the encoding strategy employed, including Pixel Encoding, Semantic Encoding, Learnable Query Encoding, and Hybrid Encoding. We illustrate the architectural variations within these categories and their corresponding encoder-decoder configurations.* #### Diffusion | Name | Title | Venue | Date | Code | Demo | | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
Excerpt of 105,411 characters
Read on GitHubWould you bet a product on this? Bounded 0–100 and slow moving.
matched fp:8c9c62dec5cc7c89, topic:vision-language-model, name:multimodal, desc:multimodal