Interleaving Reasoning for Better Text-to-Image Generation
Wenxuan Huang, Shuang Chen, Zheyong Xie, Shaosheng Cao, SHIXIANG TANG, Yufan Shen, Qingyu Yin, Wenbo Hu, Xiaoman Wang, Yuntian Tang, Junbo Qiao, Hangyu Guo
Abstract
Unified multimodal understanding and generation models recently have achieve significant improvement in image generation capability, yet a large gap remains in instruction following and detail preservation compared to systems that tightly couple comprehension with generation such as GPT-4o. Motivated by recent advances in interleaving reasoning, we explore whether such reasoning can further improve text-to-image (T2I) generation. We introduce Interleaving Reasoning Generation (IRG), a framework that alternates between text-based thinking and image synthesis: the model first produces a text-based thinking to guide an initial image, then reflects on the result to refine fine-grained details, visual quality, and aesthetics while preserving semantics. To train IRG effectively, we propose Interleaving Reasoning Generation Learning (IRGL), which targets two sub-goals: (1) strengthening the initial think-and-generate stage to establish core content and base quality, and (2) enabling high-quality textual reflection and faithful implementation of those refinements in a subsequent image. We curate IRGL-300K, a 300K-scale dataset organized into six decomposed learning modes that jointly cover learning text-based thinking, and full thinking–image trajectories. Starting from a unified foundation model that natively emits interleaved text–image outputs, our two-stage training first builds robust thinking and reflection, then efficiently tunes the IRG pipeline in the full thinking–image trajectory data. Extensive experiments show SoTA performance, yielding absolute gains of 5–10 points on GenEval, WISE, TIIF, GenAI-Bench, and OneIG-EN, alongside substantial improvements in visual quality and fine-grained fidelity. As an early exploration, our results demonstrate that interleaving reasoning is a powerful paradigm for advancing T2I. The code, model weights and datasets will be released in: https://github.com/Osilly/Interleaving-Reasoning-Generation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0cdbe75a-6ae6-4286-9ce0-73539798984eCited by top-tier papers23
- Revisual-R1: Advancing Multimodal Reasoning From Optimized Cold Start to Staged Reinforcement LearningShuang Chen, Hangyu Guo, Zhaochen Su, Yafu Li et al.ICLR 2026 · 49 citations
- Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?Ouxiang Li, Yuan Wang, Xinting Hu, Huijuan Huang et al.ICLR 2026 · 39 citations
- ReasonEdit: Towards Reasoning-Enhanced Image Editing ModelsFukun Yin, Shiyu Liu, Yucheng Han, Zhibo Wang et al.CVPR 2026 · 25 citations
- Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual GenerationZiyu Guo, Renrui Zhang, Hongyu Li, Manyuan Zhang et al.CVPR 2026 · 18 citations
- Agentic Jigsaw Interaction Learning for Enhancing Visual Perception and Reasoning in Vision-Language ModelsYu Zeng, Wenxuan Huang, Shiting Huang, Xikun Bao et al.ICLR 2026 · 11 citations
Builds on25
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- NExT-GPT: Any-to-Any Multimodal LLMShengqiong Wu, Hao Fei, Leigang Qu, Wei Ji et al.ICML 2024 · 786 citations
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language ModelsWenxuan Huang, Bohan Jia, Shaosheng Cao, Zheyu Ye et al.ICLR 2026 · 670 citations
Related papers
- GIR-Bench: Versatile Benchmark for Generating Images with ReasoningHongxiang Li, Yaowei Li, Bin Lin, Yuwei Niu et al.ICLR 2026 · 15 citations
- Think-Then-Generate: Reasoning-Aware Text-to-Image Diffusion with LLM EncodersSiqi Kou, Jiachun Jin, Zetong Zhou, YE MA et al.ICML 2026 · 13 citations
- MILR: Improving Multimodal Image Generation via Test-Time Latent ReasoningYapeng Mi, Yanpeng Zhao, Hengli Li, Chenxi Li et al.ICLR 2026 · 8 citations
- Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image EditingZiyun Zeng, David Junhao Zhang, Wei Li, Mike Zheng ShouICLR 2026 · 4 citations
- Towards Unified Multimodal Interleaved Generation via Group Relative Policy OptimizationMing Nie, Chunwei Wang, Jianhua Han, Hang Xu et al.NeurIPS 2025 · 7 citations
