Show, Don't Tell: Morphing Latent Reasoning into Image Generation
Harold Haodong Chen, Xinxiang Yin, Wenjie Shu, Hongfei (Faye) Zhang, Zixin Zhang, Chenfei Liao, Litao Guo, Qifeng Chen, YINGCONG CHEN
Abstract
Text-to-image (T2I) generation has achieved remarkable progress, yet existing methods often lack the ability to dynamically reason and refine during generation-a hallmark of human creativity. Current reasoning-augmented paradigms mostly rely on explicit thought processes, where intermediate reasoning is decoded into discrete text at fixed steps with frequent image decoding and reencoding, leading to inefficiencies, information loss, and cognitive mismatches. To bridge this gap, we introduce LatentMorph, a novel framework that seamlessly integrates implicit latent reasoning into the T2I generation process. At its core, LatentMorph introduces four lightweight components: (i) a condenser for summarizing intermediate generation states into compact visual memory, (ii) a translator for converting latent thoughts into actionable guidance, (iii) a shaper for dynamically steering next image token predictions, and (iv) an RL-trained invoker for adaptively determining when to invoke reasoning. By performing reasoning entirely in continuous latent spaces, LatentMorph avoids the bottlenecks of explicit reasoning and enables more adaptive selfrefinement. Extensive experiments demonstrate that LatentMorph (I) enhances the base model Janus-Pro by 16% on GenEval and 25% on T2I-CompBench; (II) outperforms explicit paradigms (e.g., TwiG) by 15% and 11% on abstract reasoning tasks like WISE and IPV-Txt, (III) while reducing inference time by 44% and token consumption by 51%; and (IV) exhibits 71% cognitive alignment with human intuition on reasoning invocation. Our code: LatentMorph. * Equal contribution 1 HKUST(GZ) 2 HKUST 3 NWPU 4 Knowin.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8cbdf210-2a14-4be9-a737-5f243142c2baBuilds on22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Human Preference Score: Better Aligning Text-to-image Models with Human PreferenceXiaoshi Wu, Keqiang Sun, Feng Zhu, Rui Zhao et al.ICCV 2023 · 323 citations
Related papers
- MILR: Improving Multimodal Image Generation via Test-Time Latent ReasoningYapeng Mi, Yanpeng Zhao, Hengli Li, Chenxi Li et al.ICLR 2026 · 8 citations
- Interleaving Reasoning for Better Text-to-Image GenerationWenxuan Huang, Shuang Chen, Zheyong Xie, Shaosheng Cao et al.ICLR 2026 · 40 citations
- T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoTDongzhi Jiang, Ziyu Guo, Renrui Zhang, Zhuofan Zong et al.NeurIPS 2025 · 181 citations
- ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought ReasoningJiawei Gu, Yunzhuo Hao, Huichen Will Wang, Linjie Li et al.ICLR 2026 · 51 citations
- Think-Then-Generate: Reasoning-Aware Text-to-Image Diffusion with LLM EncodersSiqi Kou, Jiachun Jin, Zetong Zhou, YE MA et al.ICML 2026 · 13 citations
