Uni-CoT: Towards Unified Chain-of-Thought Reasoning Across Text and Vision
Luozheng Qin, Jia Gong, Yuqing Sun, Tianjiao Li, Haoyu Pan, Mengping Yang, Xiaomeng Yang, Chao Qu, Zhiyu Tan, Hao Li
Abstract
Chain-of-Thought (CoT) reasoning has been widely adopted to enhance Large Language Models (LLMs) by decomposing complex tasks into simpler, sequential subtasks. However, extending CoT to vision-language reasoning tasks remains challenging, as it often requires interpreting transitions of visual states to support reasoning. Existing methods often struggle with this due to limited capacity of modeling visual state transitions or incoherent visual trajectories caused by fragmented architectures. To overcome these limitations, we propose Uni-CoT, a Unified Chain-of-Thought framework that enables coherent and grounded multimodal reasoning within a single unified model. The key idea is to leverage a model capable of both image understanding and generation to reason over visual content and model evolving visual states. However, empowering a unified model to achieve that is non-trivial, given the high computational cost and the burden of training. To address this, Uni-CoT introduces a novel two-level reasoning paradigm: A Macro-Level CoT for high-level task planning and A Micro-Level CoT for subtask execution. This design significantly reduces the computational overhead. Furthermore, we introduce a structured training paradigm that combines interleaved image-text supervision for macro-level CoT with multi-task objectives for micro-level CoT. Together, these innovations allow Uni-CoT to perform scalable and coherent multi-modal reasoning. Furthermore, thanks to our design, all experiments can be efficiently completed using only 8 A100 GPUs with 80GB VRAM each. Experimental results on reasoning-driven image generation benchmark (WISE) and editing benchmarks (RISE and KRIS) indicates that Uni-CoT demonstrates SOTA performance and strong generalization, establishing Uni-CoT as a promising solution for multi-modal reasoning. Project Page and Code: https://sais-fuxi.github.io/projects/uni-cot/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0da5a427-a9e1-4c56-ba13-a2c402b54e02Cited by top-tier papers23
- ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought ReasoningJiawei Gu, Yunzhuo Hao, Huichen Will Wang, Linjie Li et al.ICLR 2026 · 51 citations
- ReasonEdit: Towards Reasoning-Enhanced Image Editing ModelsFukun Yin, Shiyu Liu, Yucheng Han, Zhibo Wang et al.CVPR 2026 · 25 citations
- Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual GenerationZiyu Guo, Renrui Zhang, Hongyu Li, Manyuan Zhang et al.CVPR 2026 · 18 citations
- Think-Then-Generate: Reasoning-Aware Text-to-Image Diffusion with LLM EncodersSiqi Kou, Jiachun Jin, Zetong Zhou, YE MA et al.ICML 2026 · 13 citations
- ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal GenerationYongyuan Liang, Wei Chow, Feng Li, Ziqiao Ma et al.ICLR 2026 · 13 citations
Builds on33
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- Long Range Arena : A Benchmark for Efficient TransformersYi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen et al.ICLR 2021 · 881 citations
Related papers
- Unsupervised Visual Chain-of-Thought Reasoning via Preference OptimizationKesen Zhao, Beier Zhu, Qianru Sun, Hanwang ZhangICCV 2025 · 3 citations
- Vinci: Deep Thinking in Text-to-Image Generation using Unified Model with Reinforcement LearningWang Lin, Wentao Hu, Liyu Jia, Kaihang Pan et al.NeurIPS 2025
- Interleaved-Modal Chain-of-ThoughtJun Gao, Yongqi Li, Ziqiang Cao, Wenjie LiCVPR 2025
- Chain-of-Thought Guided Multi-Modal Object Re-IdentificationYa Gao, Shihao Li, Zhaojun Liu, Aihua Zheng et al.CVPR 2026
- CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language ModelsZihui Cheng, Qiguang Chen, Jin Zhang, Hao Fei et al.AAAI 2025 · 36 citations
