Look-Back: Implicit Visual Re-focusing in MLLM Reasoning
Shuo Yang, Yuwei Niu, Yuyang Liu, Yang Ye, Bin Lin, Li Yuan
摘要
Multimodal Large Language Models (MLLMs) have achieved remarkable progress in multimodal reasoning. However, they often excessively rely on textual information during the later stages of inference, neglecting the crucial integration of visual input. Current methods typically address this by explicitly injecting visual information to guide the reasoning process. In this work, through an analysis of MLLM attention patterns, we made an intriguing observation: with appropriate guidance, MLLMs can spontaneously re-focus their attention on visual inputs during the later stages of reasoning, even without explicit visual information injection. This spontaneous shift in focus suggests that MLLMs are intrinsically capable of performing visual fusion reasoning. Building on this insight, we introduce Look-Back, an implicit approach designed to guide MLLMs to "look back" at visual information in a self-directed manner during reasoning. Look-Back empowers the model to autonomously determine when, where, and how to re-focus on visual inputs, eliminating the need for explicit model-structure constraints or additional input. We demonstrate that Look-Back significantly enhances the model's reasoning and perception capabilities, as evidenced by extensive empirical evaluations on multiple multimodal benchmarks. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- More Thought, Less Accuracy? On the Dual Nature of Reasoning in Vision-Language ModelsXinyu Tian, Shu Zou, Zhaoyuan Yang, Mengqi He 等ICLR 2026 · 被引用 29 次
- Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal PerceptionLai Wei, Liangbo He, jun lan, Lingzhong Dong 等ICML 2026 · 被引用 27 次
- Perceptual-Evidence Anchored Reinforced Learning for Multimodal ReasoningChi Zhang, Haibo Qiu, Qiming Zhang, Yufei Xu 等CVPR 2026 · 被引用 22 次
- AsFT: Anchoring Safety During LLM Fine-Tuning Within Narrow Safety BasinShuo Yang, Qihui Zhang, Yuyang Liu, Yue Huang 等AAAI 2026 · 被引用 19 次
- REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video UnderstandingJiaze Li, Hao Yin, Wenhui Tan, Jingyang Chen 等CVPR 2026 · 被引用 14 次
它引用的顶会 Paper31
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual ContextsPan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu 等ICLR 2024 · 被引用 1,472 次
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language ModelsWenxuan Huang, Bohan Jia, Shaosheng Cao, Zheyu Ye 等ICLR 2026 · 被引用 670 次
- Visual-RFT: Visual Reinforcement Fine-TuningZiyu Liu, Zeyi Sun, Yuhang Zang, Xiaoyi Dong 等ICCV 2025 · 被引用 563 次
相关 Paper
- When to Think and When to Look: Uncertainty-Guided LookbackJing Bi, Filippos Bellos, JunJia Guo, Yayuan Li 等CVPR 2026 · 被引用 4 次
- Where Does Vision Meet Language? Understanding and Refining Visual Fusion in MLLMs via Contrastive AttentionShezheng Song, Shasha Li, Shan Zhao, Xiaopeng Li 等CVPR 2026 · 被引用 3 次
- Compose and Fuse: Revisiting the Foundational Bottlenecks in Multimodal ReasoningYucheng Wang, Yifan Hou, Aydin Javadov, Mubashara Akhtar 等ICLR 2026 · 被引用 3 次
- Mitigating Visual Forgetting via Take-along Visual Conditioning for Multi-modal Long CoT ReasoningHai-Long Sun, Zhun Sun, Houwen Peng, Han-Jia YeACL 2025 · 被引用 23 次
- ProReason: Multi-Modal Proactive Reasoning with Decoupled Eyesight and WisdomJingqi Zhou, Sheng Wang, Jingwei Dong, Kai Liu 等EMNLP 2025
