Reflect-DiT: Inference-Time Scaling for Text-to-Image Diffusion Transformers via In-Context Reflection
Shufan Li, Konstantinos Kallidromitis, Akash Gokul, Arsh Koneru, Yusuke Kato, Kazuki Kozuka, Aditya Grover
Abstract
The predominant approach to advancing text-to-image generation has been training-time scaling, where larger models are trained on more data using greater computational resources. While effective, this approach is computationally expensive, leading to growing interest in inference-time scaling to improve performance. Currently, inference-time scaling for text-to-image diffusion models is largely limited to best-of-N sampling, where multiple images are generated per prompt and a selection model chooses the best output. Inspired by the recent success of reasoning models like DeepSeek-R1 in the language domain, we introduce an alternative to naive best-of- sampling by equipping text-to-image Diffusion Transformers with in-context reflection capabilities. We propose Reflect-DiT, a method that enables Diffusion Transformers to refine their generations using incontext examples of previously generated images alongside textual feedback describing necessary improvements. Instead of passively relying on random sampling and hoping for a better result in a future generation, Reflect-DiT explicitly tailors its generations to address specific aspects requiring enhancement. Experimental results demonstrate that Reflect-DiT improves performance on the GenEval benchmark (+0.19) using SANA-1.0-1.6B as a base model. Additionally, it achieves a new state-of-the-art score of 0.81 on GenEval while generating only 20 samples per prompt, surpassing the previous best score of 0.80, which was obtained using a significantly larger model (SANA-1.5-4.8B) with 2048 samples under the best-of-N approach. 11Code will be available at https://github.com/jacklishufan/Reflect-DiT
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cf13141a-09a1-49c0-874c-b4091df823e0Cited by top-tier papers18
- Inference-time scaling of diffusion models through classical searchXiangcheng Zhang, Haowei Lin, Haotian Ye, James Y. Zou et al.ICLR 2026 · 57 citations
- PosterCraft: Rethinking High-Quality Aesthetic Poster Generation in a Unified FrameworkSixiang Chen, Jianyu Lai, Jialin Gao, Tian Ye et al.ICLR 2026 · 43 citations
- Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?Ouxiang Li, Yuan Wang, Xinting Hu, Huijuan Huang et al.ICLR 2026 · 39 citations
- ReasonEdit: Towards Reasoning-Enhanced Image Editing ModelsFukun Yin, Shiyu Liu, Yucheng Han, Zhibo Wang et al.CVPR 2026 · 25 citations
- Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual GenerationZiyu Guo, Renrui Zhang, Hongyu Li, Manyuan Zhang et al.CVPR 2026 · 18 citations
Builds on36
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
Related papers
- From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection TuningLe Zhuo, Liangbing Zhao, Sayak Paul, Yue Liao et al.ICCV 2025 · 5 citations
- SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion TransformerEnze Xie, Junsong Chen, Yuyang Zhao, Jincheng Yu et al.ICML 2025
- RAISE: Requirement-Adaptive Evolutionary Refinement for Training-Free Text-to-Image AlignmentLiyao Jiang, Ruichen Chen, Chao Gao, Di NiuCVPR 2026 · 7 citations
- EC-DIT: Scaling Diffusion Transformers with Adaptive Expert-Choice RoutingHaotian Sun, Tao Lei, Bowen Zhang, Yanghao Li et al.ICLR 2025
- Scaling Inference Time Compute for Diffusion ModelsNanye Ma, Shangyuan Tong, Haolin Jia, Hexiang Hu et al.CVPR 2025
