tulerfeng/Gen-Searcher
quality grade D, 44 out of 100Gen-Searcher: Reinforcing Agentic Search for Image Generation
- stars
- 377
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Diffusion models, image editing, upscaling and the surrounding creative tooling.
Signals: stable-diffusion, diffusion-models, image-generation, text-to-image, generative-art, comfyui, controlnet, flux
880 results
Gen-Searcher: Reinforcing Agentic Search for Image Generation
AI image generation CLI powered by Gemini 3 Pro. Green screen transparency, reference images, style transfer. Also a Claude Code plugin.
Claude Code skill: generate UI design mockups with GPT Image 2 via ZenMux
A Claude Code skill to generate images with Nano Banana
GPT-Image-2 驱动的电商素材一键生成 Claude Code Skill | E-commerce image generation skill powered by GPT-Image-2 via Codex CLI, with 25 built-in scene templates
Natural language → ComfyUI workflow JSON. 34 built-in templates, 360+ node definitions, auto model download. Supports txt2img, img2img, txt2vid, img2vid, audio, 3D generation across SD1.5/SDXL/SD3/FLUX/Wan2.2/HunyuanVideo/LTXV/Mochi/Cosmos + LLM integration. Works as a skill for Claude Code, Cursor, and other AI coding agents.
AudioStory: Generating Long-Form Narrative Audio with Large Language Models
Fine-tune Stable Audio Open with DiT ControlNet.
Freeze the Discriminator: a Simple Baseline for Fine-Tuning GANs (CVPRW 2020)
一个面向 claude code / Codex / OpenClaw 的跨境电商和国内电商通用视觉创作 Skill。我精选25个高质量案例,涵盖纯色底产品主图、场景化生活图、平铺图、电商详情图、真实场景等等,全部配完整提示词,都可以利用 GPT-Image-2 API生成最终效果。一键生成电商相关图片!输入产品图片和需求描述,自动生成完整的电商主图、详情页图片、社媒推广图、直播间场景图等全套视觉素材。 与众不同之处是:**Campaign Style Lock** 机制 和 **强推广** 和 **重视转化效果**。
AI-powered article illustrations with intelligent position detection and cover learning system. Claude Code Skill.
AI image generation skill for Claude Code -- Creative Director powered by Gemini
Official implementation of "Ctrl-X: Controlling Structure and Appearance for Text-To-Image Generation Without Guidance" (NeurIPS 2024)
[ICLR 2025 spotlight] 3DIS: Depth-Driven Decoupled Instance Synthesis for Text-to-Image Generation
ComfyUI-OmniGen - A ComfyUI custom node implementation of OmniGen, a powerful text-to-image generation and editing model.
Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders
ERNIE-Image is an open text-to-image generation model developed by the ERNIE-Image team at Baidu. It is built on a single-stream Diffusion Transformer (DiT), with only 8B DiT parameters, it reaches state-of-the-art performance among open-weight text-to-image models.
[ECCV 2024] Bridging Different Language Models and Generative Vision Models for Text-to-Image Generation
CosmicMan: A Text-to-Image Foundation Model for Humans (CVPR 2024)
Pytorch implementation of Make-A-Scene: Scene-Based Text-to-Image Generation with Human Priors
[ICCV 2023] BoxDiff: Text-to-Image Synthesis with Training-Free Box-Constrained Diffusion
[ECCV 2026 Oral] RealGen: Photorealistic Text-to-Image Generation via Detector-Guided Rewards.
Code for Enhancing Detail Preservation for Customized Text-to-Image Generation: A Regularization-Free Approach
Cross-Modal Implicit Relation Reasoning and Aligning for Text-to-Image Person Retrieval (CVPR 2023)
24,523 repositories in the index in total.