Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Official repository of "GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing"
| Date | Stars |
|---|---|
| 2026-07-31 | 317 |
| 2026-08-06 | 317 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing <div align="center"> <a href="https://github.com/rongyaofang/GoT"><img src="https://img.shields.io/badge/Project-Homepage-green" alt="Home"></a> <a href="https://arxiv.org/abs/2503.10639"><img src="https://img.shields.io/badge/ArXiv-2503.10639-red"></a> [Rongyao Fang](https://scholar.google.com/citations?user=FtH3CW4AAAAJ&hl=en)<sup>1\*</sup>, [Chengqi Duan](https://scholar.google.com/citations?user=r9qb4ZwAAAAJ&hl=zh-CN)<sup>2\*</sup>, [Kun Wang]()<sup>3</sup>, [Linjiang Huang](https://leonhlj.github.io/)<sup>6</sup>, [Hao Li](https://scholar.google.com/citations?user=qHqQsY4AAAAJ&hl=zh-CN)<sup>1,4</sup>, [Shilin Yan](https://scholar.google.com/citations?user=2VhjOykAAAAJ&hl=zh-CN), [Hao Tian]()<sup>3</sup>, [Xingyu Zeng]()<sup>3</sup>, [Rui Zhao]()<sup>3</sup>, [Jifeng Dai](https://jifengdai.org/)<sup>4,5</sup>, [Xihui Liu](https://xh-liu.github.io/)<sup>2 :envelope:</sup>, [Hongsheng Li](https://www.ee.cuhk.edu.hk/~hsli/)<sup>1 :envelope:</sup> <sup>1</sup>CUHK MMLab, <sup>2</sup>HKU MMLab, <sup>3</sup>SenseTime, <sup>4</sup>Shanghai AI Laboratory, <sup>5</sup>Tsinghua University, <sup>6</sup>Beihang University *Equal contribution, :envelope:Corresponding authors </div> <div align="center"> <img src="figures/teaser.jpg" width="100%" alt="GoT Framework" /> </div> <hr> <div align="center" style="line-height: 1.2;"> <a href="https://arxiv.org/abs/2503.10639" target="_blank"><b>Paper</b></a> • <a href="#introduction">Introduction</a> • <a href="#released-datasets">Datasets</a> • <a href="#released-model-got-framework">Model</a> • <a href="#results">Results</a> • <a href="https://huggingface.co/LucasFang/GoT-6B" target="_blank">🤗 Hugging Face</a> • <a href="#license">License</a> </div> ## 🔥 News - **[2025-9-19]** 📝 Our GoT paper has been accepted by **NeurIPS 2025**! - **[2025-9-12]** 🎉 We open-sourced our latest work **FLUX-Reason-6M** dataset! This high-quality text-to-image reasoning dataset was constructed using 15,000 A100 GPU days with FLUX generation. Check it out at [FLUX-Reason-6M](https://github.com/rongyaofang/prism-bench)! ## Introduction We present **Generation Chain-of-Thought (GoT)**, a novel paradigm that enables generation and editing through an explicit language reasoning process before outputting images. This approach transforms conventional text-to-image generation and editing into a reasoning-guided framework that analyzes semantic relationships and spatial arrangements. GoT pioneers a new direction for reasoning-driven visual generation and editing, producing images that better align with human intent through: - **Semantic-Spatial Reasoning**: Integrates both semantic understanding and explicit spatial coordinates - **Unified Framework**: Handles both image generation and editing with the same architecture ## Released Datasets | Dataset | Link | Amount | |---------|------|--------| | **Laion-Aesthetics-High-Resolution-GoT** | [🤗 HuggingFace](https://huggingface.co/datasets/LucasFang/Laion-Aesthetics-High-Resolution-GoT) | 3.77M | | **JourneyDB-GoT** | [🤗 HuggingFace](https://huggingface.co/datasets/LucasFang/JourneyDB-GoT) | 4.09M | | **OmniEdit-GoT** | [🤗 HuggingFace](https://huggingface.co/datasets/LucasFang/OmniEdit-GoT) | 736K | | **FLUX-Reason-6M** | [🤗 HuggingFace](https://huggingface.co/datasets/LucasFang/FLUX-Reason-6M) | 6M | ## Dataset Features ### Laion-Aesthetics-High-Resolution-GoT - 3.77 million High-quality images filtered for sizes larger than 512 pixels from Laion-Aesthetics - Prompts and GoT descriptions from Qwen2-VL - Prompts averaging 110.81 characters - GoT descriptions averaging 811.56 characters - 3.78 bounding boxes per image on average ### JourneyDB-GoT - 4.09 million high-quality AI-generated images - Prompts and GoT descriptions from Qwen2-VL - Prompts averaging 149.78 characters - GoT descriptions averaging 906.01 characters -
Excerpt of 8,947 characters
Read on GitHubWould you bet a product on this? Bounded 0–100 and slow moving.
matched fp:d689d6b6f8507893, desc:multimodal