Diffusion Guided Chain-of-Vision for Large Autoregressive Vision Models
Xinyang Wang, Kecheng Zheng, Minfeng Zhu, Wei Wu, Fan Lu, Wei Zhai, Wei Chen
Abstract
Chain-of-Thought (CoT) has recently shown encouraging progress in the vision language model. However, the pure-vision CoT (i.e., chain-of-vision) has been underexplored in visual in-context learning. In this paper, we introduce Diffusion Guided Chain-of-Vision, which integrates an explicit chain-of-thought process into autoregressive vision models through vision prior from pre-trained diffusion models. Concretely, we find that pre-trained diffusion models induce a reliable probability flow in image space, where intermediate images sampled along this flow exhibit visual coherence and serve as task-free, chain-of-vision supervision for pure-vision autoregressive models. Extensive experiments on diverse vision tasks and multi-scale models validate the effectiveness of our proposed method for visual in-context learning. Code and dataset will be publicly available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on46
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- In-Context Learning Unlocked for Diffusion ModelsZhendong Wang, Yifan Jiang, Yadong Lu, Yelong Shen et al.NeurIPS 2023 · 128 citations
- Chain-of-Focus Prompting: Leveraging Sequential Visual Cues to Prompt Large Autoregressive Vision ModelsJiyang Zheng, Jialiang Shen, Yu Yao, Min Wang et al.ICLR 2025
- CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action ModelsQingqing Zhao, Yao Lu, Moo Jin Kim, Zipeng Fu et al.CVPR 2025
- Multi-Modal Latent Space Learning for Chain-of-Thought Reasoning in Language ModelsLiqi He, Zuchao Li, Xiantao Cai, Ping WangAAAI 2024 · 38 citations
- Interleaved-Modal Chain-of-ThoughtJun Gao, Yongqi Li, Ziqiang Cao, Wenjie LiCVPR 2025
