From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection Tuning
Le Zhuo, Liangbing Zhao, Sayak Paul, Yue Liao, Renrui Zhang, Yi Xin, Peng Gao, Mohamed Elhoseiny, Hongsheng Li
Abstract
Recent text-to-image diffusion models achieve impressive visual quality through extensive scaling of training data and model parameters, yet they often struggle with complex scenes and fine-grained details. Inspired by the selfreflection capabilities emergent in large language models, we propose ReflectionFlow, an inference-time framework enabling diffusion models to iteratively reflect upon and refine their outputs. ReflectionFlow introduces three complementary inference-time scaling axes: (1) noise-level scaling to optimize latent initialization; (2) prompt-level scaling for precise semantic guidance; and most notably, (3) reflection-level scaling, which explicitly provides actionable reflections to iteratively assess and correct previous generations. To facilitate reflection-level scaling, we construct GenRef, a large-scale dataset comprising 1 million triplets, each containing a reflection, a flawed image, and an en- * Equal Contribution † Corresponding Author hanced image. Leveraging this dataset, we efficiently perform reflection tuning on state-of-the-art diffusion transformer, FLUX.1-dev, by jointly modeling multimodal inputs within a unified framework. Experimental results show that ReflectionFlow significantly outperforms naive noise-level scaling methods, offering a scalable and compute-efficient solution toward higher-quality image synthesis on challenging tasks. All code, checkpoints, and datasets are available at https://diffusion-cot.github.io/ reflection2perfection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d0f9246b-14ea-44b5-affa-e3137248b0baCited by top-tier papers17
- MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought ReasoningXinyan Chen, Renrui Zhang, Dongzhi Jiang, Aojun Zhou et al.NeurIPS 2025 · 54 citations
- Interleaving Reasoning for Better Text-to-Image GenerationWenxuan Huang, Shuang Chen, Zheyong Xie, Shaosheng Cao et al.ICLR 2026 · 40 citations
- Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual GenerationZiyu Guo, Renrui Zhang, Hongyu Li, Manyuan Zhang et al.CVPR 2026 · 18 citations
- Factuality Matters: When Image Generation and Editing Meet Structured VisualsLe Zhuo, Songhao Han, Yuandong Pu, Boxiang Qiu et al.ICLR 2026 · 15 citations
- Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and GenerationShufan Li, Jiuxiang Gu, Kangning Liu, Zhe Lin et al.ICLR 2026 · 14 citations
Builds on41
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
Related papers
- Reflect-DiT: Inference-Time Scaling for Text-to-Image Diffusion Transformers via In-Context ReflectionShufan Li, Konstantinos Kallidromitis, Akash Gokul, Arsh Koneru et al.ICCV 2025 · 3 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- Zigzag Diffusion Sampling: Diffusion Models Can Self-Improve via Self-ReflectionLichen Bai, Shitong Shao, Zikai Zhou, Zipeng Qi et al.ICLR 2025
- Scaling Inference Time Compute for Diffusion ModelsNanye Ma, Shangyuan Tong, Haolin Jia, Hexiang Hu et al.CVPR 2025
- MirrorVerse: Pushing Diffusion Models to Realistically Reflect the WorldAnkit Dhiman, Manan Shah, R. Venkatesh BabuCVPR 2025
