Mitigating Visual Forgetting via Take-along Visual Conditioning for Multi-modal Long CoT Reasoning
Hai-Long Sun, Zhun Sun, Houwen Peng, Han-Jia Ye
摘要
Recent advancements in Large Language Models (LLMs) have demonstrated enhanced reasoning capabilities, evolving from Chainof-Thought (CoT) prompting to advanced, product-oriented solutions like OpenAI o1. During our re-implementation of this model, we noticed that in multimodal tasks requiring visual input (e.g., geometry problems), Multimodal LLMs (MLLMs) struggle to maintain focus on the visual information, in other words, MLLMs suffer from a gradual decline in attention to visual information as reasoning progresses, causing text-over-relied outputs. To investigate this, we ablate image inputs during long-chain reasoning. Concretely, we truncate the reasoning process midway, then re-complete the reasoning process with the input image removed. We observe only a ∼2% accuracy drop on MathVista's test-hard subset, revealing the model's textual outputs dominate the following reasoning process. Motivated by this, we propose Take-along Visual Conditioning (TVC), a strategy that shifts image input to critical reasoning stages and compresses redundant visual tokens via dynamic pruning. This methodology helps the model retain attention to the visual components throughout the reasoning. Our approach achieves state-of-the-art performance on average across five mathematical reasoning benchmarks (+3.4 points vs previous sota), demonstrating the effectiveness of TVC in enhancing multimodal reasoning systems. The project page is available at https: //sun-hailong.github.io/projects/TVC .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware ReasoningSARA GHAZANFARI, Francesco Croce, Nicolas Flammarion, Prashanth Krishnamurthy 等CVPR 2026 · 被引用 35 次
- VisMem: Latent Vision Memory Unlocks Potential of Vision-Language ModelsXinlei Yu, Chengming Xu, Guibin Zhang, Zhangquan Chen 等CVPR 2026 · 被引用 30 次
- Multimodal Tabular Reasoning with Privileged Structured InformationJun-Peng Jiang, Yu Xia, Hai-Long Sun, Shiyin Lu 等NeurIPS 2025 · 被引用 16 次
- Triplets Better Than Pairs: Towards Stable and Effective Self-Play Fine-Tuning for LLMsYibo Wang, Hai-Long Sun, Guangda Huzhang, Qingguo Chen 等NeurIPS 2025 · 被引用 12 次
- Visually-Guided Policy Optimization for Multimodal ReasoningZengbin Wang, Feng Xiong, Liang Lin, Xuecai Hu 等ACL 2026 · 被引用 7 次
它引用的顶会 Paper15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language ModelsYushi Hu, Weijia Shi, Xingyu Fu, Dan Roth 等NeurIPS 2024 · 被引用 373 次
相关 Paper
- TVI-CoT: Text-Visual Interleaved Chain-of-Thought Reasoning for Multimodal UnderstandingLianyu Hu, Xiaoyu Ma, Zeqin Liao, Yang LiuICML 2026 · 被引用 2 次
- See It, Say It, Sorted: An Iterative Training-Free Framework for Visually-Grounded Multimodal Reasoning in LVLMsYongchang Zhang, Xianzheng Ma, Tianyi Liu, Guangquan Zhou 等CVPR 2026 · 被引用 2 次
- ProReason: Multi-Modal Proactive Reasoning with Decoupled Eyesight and WisdomJingqi Zhou, Sheng Wang, Jingwei Dong, Kai Liu 等EMNLP 2025
- MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought ReasoningXinyan Chen, Renrui Zhang, Dongzhi Jiang, Aojun Zhou 等NeurIPS 2025 · 被引用 54 次
- Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language ModelsYuhao Dong, Zuyan Liu, Hai-Long Sun, Jingkang Yang 等CVPR 2025
