ZitengWangNYU/Scale-RAE
quality grade D, 43 out of 100Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders
- stars
- 263
- stars gained this week
- —this week
- forks, open issues and contributors
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Diffusion models, image editing, upscaling and the surrounding creative tooling.
Signals: stable-diffusion, diffusion-models, image-generation, text-to-image, generative-art, comfyui, controlnet, flux
879 results
Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders
ERNIE-Image is an open text-to-image generation model developed by the ERNIE-Image team at Baidu. It is built on a single-stream Diffusion Transformer (DiT), with only 8B DiT parameters, it reaches state-of-the-art performance among open-weight text-to-image models.
[ECCV 2024] Bridging Different Language Models and Generative Vision Models for Text-to-Image Generation
CosmicMan: A Text-to-Image Foundation Model for Humans (CVPR 2024)
Pytorch implementation of Make-A-Scene: Scene-Based Text-to-Image Generation with Human Priors
[ICCV 2023] BoxDiff: Text-to-Image Synthesis with Training-Free Box-Constrained Diffusion
[ECCV 2026 Oral] RealGen: Photorealistic Text-to-Image Generation via Detector-Guided Rewards.
Code for Enhancing Detail Preservation for Customized Text-to-Image Generation: A Regularization-Free Approach
Cross-Modal Implicit Relation Reasoning and Aligning for Text-to-Image Person Retrieval (CVPR 2023)
Official implementation of "Controlling Text-to-Image Diffusion by Orthogonal Finetuning".
[Tutorial] Few-Step Distillation for Text-to-Image Generation: A Practical Guide
Training-Free Structured Diffusion Guidance for Compositional Text-to-Image Synthesis
CSGO: Content-Style Composition in Text-to-Image Generation 🔥
Mini-DALLE3: Interactive Text to Image by Prompting Large Language Models
[SIGGRAPH Asia 2024, Best Paper Honorable Mention] This is the official implementation of our SIGGRAPH Asia journal artical: TEXGen: a Generative Diffusion Model for Mesh Textures
An open-source toolbox for fast sampling of diffusion models. Official implementations of our works published in ICML'24, NeurIPS'24, CVPR'24, J. Stat. Mech'25.
[CVPR2024 (Highlight)] RichDreamer: A Generalizable Normal-Depth Diffusion Model for Detail Richness in Text-to-3D. Live Demo:https://modelscope.cn/studios/Damo_XR_Lab/3D_AIGC
PyTorch implementation of the ICCV paper "3D-aware Image Generation using 2D Diffusion Models"
Official PyTorch implementation of the NeurIPS 2022 paper "Improving Diffusion Models for Inverse Problems using Manifold Constraints (https://arxiv.org/abs/2206.00941)"
From baby GPT to diffusion GPT: An annotated implementation of a character-level discrete diffusion model (adapted from Karpathy’s baby GPT).
This is the code for the paper "RadioDiff: An Effective Generative Diffusion Model for Sampling-Free Dynamic Radio Map Construction", IEEE TCCN.
Official implementation of CVPR 2024 paper: "FreeControl: Training-Free Spatial Control of Any Text-to-Image Diffusion Model with Any Condition"
[NIPS 2023] Official implementation for "DiffSketcher: Text Guided Vector Sketch Synthesis through Latent Diffusion Models" https://arxiv.org/abs/2306.14685
[ICCV 2023] Official PyTorch implementation for the paper "FreeDoM: Training-Free Energy-Guided Conditional Diffusion Model"
24,535 repositories in the index in total.