Latent Chain-of-Thought for Visual Reasoning
Guohao Sun, Hang Hua, Jian Wang, Jiebo Luo, Sohail A. Dianat, Majid Rabbani, Raghuveer Rao, Zhiqiang Tao
Abstract
Chain-of-thought (CoT) reasoning is critical for improving the interpretability and reliability of Large Vision-Language Models (LVLMs). However, existing training algorithms such as SFT, PPO, and GRPO may not generalize well across unseen reasoning tasks and heavily rely on a biased reward model. To address this challenge, we reformulate reasoning in LVLMs as posterior inference and propose a scalable training algorithm based on amortized variational inference. By leveraging diversity-seeking reinforcement learning algorithms, we introduce a novel sparse reward function for token-level learning signals that encourage diverse, high-likelihood latent CoT, overcoming deterministic sampling limitations and avoiding reward hacking. Additionally, we implement a Bayesian inference-scaling strategy that replaces costly Best-of-N and Beam Search with a marginal likelihood to efficiently rank optimal rationales and answers. We empirically demonstrate that the proposed method enhances the state-of-the-art LVLMs on seven reasoning benchmarks, in terms of effectiveness, generalization, and interpretability. The code is available at https://github.com/heliossun/LaCoT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 504cc9b3-3870-44a7-b5b1-ef12987fbadfCited by top-tier papers14
- VisMem: Latent Vision Memory Unlocks Potential of Vision-Language ModelsXinlei Yu, Chengming Xu, Guibin Zhang, Zhangquan Chen et al.CVPR 2026 · 30 citations
- Thinking in Uncertainty: Mitigating Hallucinations in MLRMs with Latent Entropy-Aware DecodingZhongxing Xu, Zhonghua Wang, Zhe Qian, Dachuan Shi et al.CVPR 2026 · 16 citations
- Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action ModelsShuanghao Bai, Jing Lyu, Wanqi Zhou, Zhe Li et al.ICML 2026 · 15 citations
- Forest Before Trees: Latent Superposition for Efficient Visual ReasoningYubo Wang, Juntian Zhang, Yichen Wu, Yankai Lin et al.ACL 2026 · 12 citations
- Latent Chain-of-Thought World Modeling for End-to-End Autonomous DrivingShuhan Tan, Kashyap Chitta, Yuxiao Chen, Ran Tian et al.CVPR 2026 · 10 citations
Builds on24
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- MM-Vet: Evaluating Large Multimodal Models for Integrated CapabilitiesWeihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang et al.ICML 2024 · 1,191 citations
- Flow Network based Generative Models for Non-Iterative Diverse Candidate GenerationEmmanuel Bengio, Moksh Jain, Maksym Korablyov, Doina Precup et al.NeurIPS 2021 · 565 citations
- Defining and Characterizing Reward GamingJoar Skalse, Nikolaus H. R. Howe, Dmitrii Krasheninnikov, David KruegerNeurIPS 2022 · 466 citations
- MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding BenchmarkXiang Yue, Tianyu Zheng, Yuansheng Ni, Yubo Wang et al.ACL 2025 · 377 citations
Related papers
- Improve Vision Language Model Chain-of-thought ReasoningRuohong Zhang, Bowen Zhang, Yanghao Li, Haotian Zhang et al.ACL 2025 · 135 citations
- ReaGEN: Adaptive Generation of Structured Chains-of-Thought for Efficient Multimodal ReasoningRuiqing Tian, Mohan Sai Singamsetti, Di Niu, Bahador RashidiCVPR 2026
- Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement LearningSimon Zhai, Hao Bai, Zipeng Lin, Jiayi Pan et al.NeurIPS 2024 · 214 citations
- The Surprising Effectiveness of Negative Reinforcement in LLM ReasoningXinyu Zhu, Mengzhou Xia, Zhepei Wei, Wei-Lin Chen et al.NeurIPS 2025 · 177 citations
- Unveiling Chain of Step Reasoning for Vision-Language Models with Fine-grained RewardsHonghao Chen, Xingzhou Lou, Xiaokun Feng, Kaiqi Huang et al.NeurIPS 2025 · 7 citations
