Does Reasoning Improve Seeing? Understanding When Vision-Language Models Benefit from Thinking
Jing Bi, Luchuan Song, Dingxin Zhang, Pinxin Liu, Guangyu Sun, Lianggong Wen, Weidong Cai, Chen Chen, Chenliang Xu
摘要
Vision-language models (VLMs) now support both direct Instruct and explicit-reasoning Thinking modes, yet practitioners still lack principled ways to decide when reasoning actually improves performance, or how much computation to spend at test time, so we investigate whether VLMs encode meta-cognitive signals for adaptive inference. We derive oracle labels for two properties: (1) reasoning helpfulness, namely whether explicit reasoning improves accuracy, and (2) desired generation length, the minimal token budget needed for a correct answer. Probing final-layer representations in InternVL and Qwen3-VL models, we find Thinking models encode these signals more linearly than Instruct models, suggesting that reasoning-oriented training enhances meta-cognitive structure. Head-wise attribution reveals two circuits: length-control heads in lower layers and reasoning/difficulty heads in higher layers. Causal interventions confirm these roles across benchmarks: scaling length heads controls output length with minimal accuracy loss, while scaling reasoning heads enables a perception-reasoning trade-off, improving accuracy by up to 5.3%. Our results demonstrate that reasoning-tuned VLMs develop localized, manipulable circuits for meta-cognitive control, enabling test-time steering of computation and reasoning without retraining.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper22
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu 等NeurIPS 2022 · 被引用 2,727 次
- Graph of Thoughts: Solving Elaborate Problems with Large Language ModelsMaciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger 等AAAI 2024 · 被引用 1,292 次
相关 Paper
- When to Think and When to Look: Uncertainty-Guided LookbackJing Bi, Filippos Bellos, JunJia Guo, Yayuan Li 等CVPR 2026 · 被引用 4 次
- ReaGEN: Adaptive Generation of Structured Chains-of-Thought for Efficient Multimodal ReasoningRuiqing Tian, Mohan Sai Singamsetti, Di Niu, Bahador RashidiCVPR 2026
- SPARC: Separating Perception And Reasoning Circuits for Test-time Scaling of VLMsNiccolò Avogaro, Nayanika Debnath, Li Mi, Thomas Frick 等ICML 2026
- ERGO: Efficient High-Resolution Visual Understanding for Vision-Language ModelsJewon Lee, Wooksu Shin, Seungmin Yang, Ki-Ung Song 等ICLR 2026 · 被引用 3 次
- Think Deep, Not Just Long: Measuring LLM Reasoning Effort via Deep-Thinking TokensWei-Lin Chen, Liqian Peng, Tian Tan, Chao Zhao 等ICML 2026 · 被引用 20 次
