When to Think and When to Look: Uncertainty-Guided Lookback
Jing Bi, Filippos Bellos, JunJia Guo, Yayuan Li, Chao Huang, Yolo Yunlong Tang, Luchuan Song, Susan Liang, Zhongfei Zhang, Jason J. Corso, Chenliang Xu
摘要
Test-time “thinking” (i.e., generating explicit intermediate reasoning chains) is known to boost performance in large language models and has recently shown strong gains for large vision–language models (LVLMs). However, despite these promising results, there is still no systematic analysis of how thinking actually affects visual reasoning. We provide the first such analysis with a large-scale, controlled comparison of thinking for LVLMs, evaluating 10 variants from the InternVL3.5 and Qwen3-VL families on MMMUval under generous token budgets and multi-pass decoding. We show that more thinking is not always better: long chains often yield long-wrong trajectories that ignore the image and underperform the same models run in standard instruct mode. A deeper analysis reveals that certain short “lookback” phrases, which explicitly refer back to the image, are strongly enriched in successful trajectories and correlate with better visual grounding. Building on this insight, we propose uncertainty-guided lookback, a training-free decoding strategy that combines an uncertainty signal with adaptive lookback prompts and breadth search. Our method improves overall MMMU performance, delivers the largest gains in categories where standard thinking is weak, and outperforms several strong decoding baselines, setting a new state of the art under fixed model families and token budgets. We further show that this decoding strategy generalizes, yielding consistent improvements on five additional benchmarks, including two broad multimodal suites and math-focused visual reasoning datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Mistake Attribution: Fine-Grained Mistake Understanding in Egocentric VideosYayuan Li, Aadit Jain, Filippos Bellos, Jason J. CorsoCVPR 2026 · 被引用 3 次
- Does Reasoning Improve Seeing? Understanding When Vision-Language Models Benefit from ThinkingJing Bi, Luchuan Song, Dingxin Zhang, Pinxin Liu 等ICML 2026
它引用的顶会 Paper18
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual ContextsPan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu 等ICLR 2024 · 被引用 1,472 次
- Are We on the Right Way for Evaluating Large Vision-Language Models?Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang 等NeurIPS 2024 · 被引用 1,029 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language ModelsYushi Hu, Weijia Shi, Xingyu Fu, Dan Roth 等NeurIPS 2024 · 被引用 373 次
- Dynamic Early Exit in Reasoning ModelsChenxu Yang, Qingyi Si, Yongjie Duan, Zheliang Zhu 等ICLR 2026 · 被引用 250 次
相关 Paper
- See It, Say It, Sorted: An Iterative Training-Free Framework for Visually-Grounded Multimodal Reasoning in LVLMsYongchang Zhang, Xianzheng Ma, Tianyi Liu, Guangquan Zhou 等CVPR 2026 · 被引用 2 次
- TVI-CoT: Text-Visual Interleaved Chain-of-Thought Reasoning for Multimodal UnderstandingLianyu Hu, Xiaoyu Ma, Zeqin Liao, Yang LiuICML 2026 · 被引用 2 次
- Rationale-Enhanced Decoding for Multi-modal Chain-of-ThoughtShin'ya Yamaguchi, Kosuke Nishida, Daiki ChijiwaCVPR 2026
- DeepScan: A Training-Free Framework for Visually Grounded Reasoning in Large Vision-Language ModelsYangfu Li, Hongjian Zhan, Jiawei Chen, Yuning Gong 等CVPR 2026 · 被引用 6 次
- Mitigating Low-Quality Reasoning in MLLMs: Self-Driven Refined Multimodal CoT with Selective Thinking and Step-wise Visual EnhancementChongjun Tu, Peng Ye, Dongzhan Zhou, Tao Chen 等AAAI 2026
