Self-Correcting LLM-Controlled Diffusion Models
Tsung-Han Wu, Long Lian, Joseph E. Gonzalez, Boyi Li, Trevor Darrell
Abstract
Text-to-image generation has witnessed significant progress with the advent of diffusion models. Despite the ability to generate photorealistic images, current text-to-image diffusion models still often struggle to accurately interpret and follow complex input text prompts. In contrast to existing models that aim to generate images only with their best effort, we introduce Self-correcting LLM-controlled Diffusion (SLD). SLD is a framework that generates an image from the input prompt, assesses its alignment with the prompt, and performs self-corrections on the inaccuracies in the generated image. Steered by an LLM controller, SLD turns text-to-image generation into an iterative closed-loop process, ensuring correctness in the resulting image. SLD is not only training-free but can also be seamlessly integrated with diffusion models behind API access, such as DALL-E 3, to further boost the performance of state-of-the-art diffusion models. Experimental results show that our approach can rectify a majority of incorrect generations, particularly in generative numeracy, attribute binding, and spatial relationships. Furthermore, by simply adjusting the instructions to the LLM, SLD can perform image editing tasks, bridging the gap between text-to-image generation and image editing pipelines. Our code is available at: https://self-correcting-llm-diffusion.github.io.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2bdadf66-5451-4e43-bfe0-166322787ff5Cited by top-tier papers49
- Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMsLing Yang, Zhaochen Yu, Chenlin Meng, Minkai Xu et al.ICML 2024 · 231 citations
- GenArtist: Multimodal LLM as an Agent for Unified Image Generation and EditingZhenyu Wang, Aoxue Li, Zhenguo Li, Xihui LiuNeurIPS 2024 · 162 citations
- SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D PriorsChenyang Ma, Kai Lu, Ta Ying Cheng, Niki Trigoni et al.NeurIPS 2024 · 82 citations
- CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept MatchingDongzhi Jiang, Guanglu Song, Xiaoshi Wu, Renrui Zhang et al.NeurIPS 2024 · 75 citations
- HiCo: Hierarchical Controllable Diffusion Model for Layout-to-image GenerationBo Cheng, Yuhang Ma, Liebucha Wu, Shanyuan Liu et al.NeurIPS 2024 · 53 citations
Builds on19
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- SDEdit: Guided Image Synthesis and Editing with Stochastic Differential EquationsChenlin Meng, Yutong He, Yang Song, Jiaming Song et al.ICLR 2022 · 2,128 citations
- MultiDiffusion: Fusing Diffusion Paths for Controlled Image GenerationOmer Bar-Tal, Lior Yariv, Yaron Lipman, Tali DekelICML 2023 · 575 citations
Related papers
- Self-Corrected Image Generation with Explainable Latent RewardsYinyi Luo, Hrishikesh Gokhale, Marios Savvides, Jindong Wang et al.CVPR 2026 · 1 citation
- Exploring the Role of Large Language Models in Prompt Encoding for Diffusion ModelsBingqi Ma, Zhuofan Zong, Guanglu Song, Hongsheng Li et al.NeurIPS 2024 · 57 citations
- LLM4GEN: Leveraging Semantic Representation of LLMs for Text-to-Image GenerationMushui Liu, Yuhang Ma, Zhen Yang, Jun Dan et al.AAAI 2025 · 36 citations
- Self-Correcting Decoding with Generative Feedback for Mitigating Hallucinations in Large Vision-Language ModelsCe Zhang, Zifu Wan, Zhehan Kan, Martin Q. Ma et al.ICLR 2025
- UniEdit-I: Training-free Image Editing for Unified VLM via Iterative Understanding, Editing and VerifyingChengyu Bai, Jintao Chen, Xiang Bai, Yilong Chen et al.CVPR 2026 · 8 citations
