Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Reading notes about Multimodal Large Language Models, Large Language Models, and Diffusion Models
| Date | Stars |
|---|---|
| 2026-07-31 | 1187 |
| 2026-08-02 | 1188 |
| 2026-08-06 | 1188 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# Awesome-Multimodal-Large-Language-Models This is a repository for organizing articles related to Multimodal Large Language Models, Large Language Models, and Diffusion Models; Most papers are linked to **my reading notes**. Feel free to visit my [personal homepage](https://yfzhang114.github.io/) and contact me for collaboration and discussion. ### About Me :high_brightness: I'm a final-year Ph.D. student at the State Key Laboratory of Pattern Recognition, the University of Chinese Academy of Sciences, advised by Prof. [Tieniu Tan](http://people.ucas.ac.cn/~tantieniu). I have also spent time at Microsoft, advised by Prof. [Jingdong Wang](https://jingdongwang2017.github.io/), alibaba DAMO Academy, work with Prof. [Rong Jin](https://scholar.google.com/citations?user=CS5uNscAAAAJ&hl=zh-CN). ### 🔥 Updated 2026-07-14 - [2026-07-14] Updated with several recent RL/Agentic RL, MLLM studies, along with their reading notes. - We present [Skywork-R1V4: Toward Agentic Multimodal Intelligence through Interleaved Thinking with Images and DeepResearch](https://skywork-r1v4-lite.netlify.app/) [[Reading Notes]](https://zhuanlan.zhihu.com/p/1979848119471608282), Skywork-R1V4 requires only 30K SFT data and activates "think with image," search, planning, and interleaved image manipulation/search capabilities, with 3B activated parameters, outperforming Gemini 2.5 Flash on all perception and deep research benchmarks. - We present [Thyme: Think Beyond Images](https://thyme-vl.github.io/) [[Reading Notes]](https://zhuanlan.zhihu.com/p/1942175827547649963), Thyme transcends traditional ``thinking with images'' paradigms by autonomously generating and executing diverse image processing and computational operations through executable code. - We present [R1-Reward](https://github.com/yfzhang114/r1_reward) [[Reading Notes]](https://zhuanlan.zhihu.com/p/1903095194166997749), which is a comprehensive project focused on enhancing multimodal reward modeling through reinforcement learning. - We present [MME-Unify](https://mme-unify.github.io/), a comprehensive benchmark for unified multimodal models (GPT-4o, Gemini-2-flash, Janus-Pro, EMU3, Show-o, VILA-U). - We present [MM-RLHF](https://github.com/yfzhang114/MM-RLHF), a comprehensive dataset of 120K fully human-annotated preference data, along with a robust reward model and training algorithm, designed to enhance MLLM alignment and significantly improve performance across 27 benchmark tasks. - Our benchmark [MME-RealWorld](https://mme-realworld.github.io/) has been released, the most difficult and largest pure manual annotation image perception benchmark so far. [[Code]](https://github.com/yfzhang114/MME-RealWorld) [[Reading Notes]](https://zhuanlan.zhihu.com/p/717129017) - Our model [SliME](https://arxiv.org/abs/2406.08487) has been released, a high-resolution MLLM that can also be extend to video analysis. [[Code]](https://github.com/yfzhang114/SliME) [[Reading Notes]](https://zhuanlan.zhihu.com/p/703258020) - Our paper [Debiasing Multimodal Large Language Models](https://arxiv.org/abs/2403.05262) has been released. [[Code]](https://github.com/yfzhang114/LLaVA-Align) [[Reading Notes]](https://zhuanlan.zhihu.com/p/686461442) # Table of Contents (ongoing) - [Awesome-Multimodal-Large-Language-Models](#awesome-multimodal-large-language-models) - [Table of Contents (ongoing)](#table-of-contents-ongoing) - [Survey and Outlook](#survey-and-outlook) - [Multimodal Reasoning & Think with Images (o3)](#multimodal-reasoning-and-think-with-images-o3) - [Multimodal Large Language Models](#multimodal-large-language-models) - [BenchMark and Dataset](#benchmark-and-dataset) - [Unify Multimodal Understanding and Generation](#unify-multimodal-understanding-and-generation) - [Alignment With Human Preference (MLLM)](#alignment-with-human-preference-mllm) - [Alignment With Human Preference (LLM)](#alignment-with-human-preference-llm) - [Training Frameworks](#training-frameworks) # Survey and Outlook 1. [GLM/Qw
Excerpt of 31,053 characters
Read on GitHubYi-Fan Zhang · State Key Laboratory of Pattern Recognition
53
20
Yue Ding · State Key Laboratory of Pattern Recognition
2
Samit · HKUST · Hong Kong
1
Wenqi Zhang · Zhejiang University · China
1
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:972475c4d815f798, name:multimodal, desc:multimodal