Reason with Thumbnails, Answer with Focus: An Efficient and Effective Paradigm for Multimodal Grounded Visual Reasoning
An-Lan Wang, Guozhi Tang, Lei Liao, Hanshen Zhu, Kai Huang, Jingqun Tang, Jiaming Zhou, Kun-Yu Lin
摘要
To enhance the interpretability of multimodal large language models' outputs, recent efforts explored Grounded Visual Reasoning (GVR), in which the model is trained to select relevant image regions before answering the question. However, the multi-round ``ground-then-answer'' and reasoning nature of these methods imposes much more computational costs compared to non-GVR methods. To attain efficient and effective GVR, in this paper, we propose a novel paradigm called Reason with Thumbnails, Answer with Focus (RTAF), which feeds the model with low-resolution images to reason the relevant regions and high-resolution crops to answer the final answer. Our motivation arises from the observation that, in many cases, the key area required to answer questions can be inferred from the low-resolution thumbnails, without the need for a full-resolution image. Additionally, for extreme cases where thumbnails lack sufficient information (leading to undirected region guessing and increased computation), we equip the model with a tool to access higher-resolution images. For training efficiency, we adopt pure reinforcement learning (i.e., GRPO) and design a suite of reward functions to supervise the model's behavior, alongside a resolution-aware training data selection strategy. Finally, our model, based on Qwen2.5-VL, achieves significant improvements across a range of benchmarks with reduced computation, demonstrating the effectiveness and efficiency of our proposed RTAF, e.g., compared to the non-GVR model Qwen2.5-VL, our model achieves a performance gain of 5.8 while using comparable visual tokens (471 vs. 391). Against state-of-the-art GVR methods, RTAF reduces visual token usage by half while delivering superior performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language ModelsWenxuan Huang, Bohan Jia, Shaosheng Cao, Zheyu Ye 等ICLR 2026 · 被引用 670 次
- DeepEyes: Incentivizing "Thinking with Images" via Reinforcement LearningZiwei Zheng, Michael Yang, Jack Hong, Chenxiao Zhao 等ICLR 2026 · 被引用 321 次
- Reward Model Ensembles Help Mitigate OveroptimizationThomas Coste, Usman Anwar, Robert Kirk, David KruegerICLR 2024 · 被引用 208 次
- SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language ModelsDongyang Liu, Renrui Zhang, Longtian Qiu, Siyuan Huang 等ICML 2024 · 被引用 149 次
相关 Paper
- ERGO: Efficient High-Resolution Visual Understanding for Vision-Language ModelsJewon Lee, Wooksu Shin, Seungmin Yang, Ki-Ung Song 等ICLR 2026 · 被引用 3 次
- Deeper Thought, Weaker Aim: Understanding and Mitigating Perceptual Impairment during Reasoning in Multimodal Large Language ModelsRuiying Peng, Xueyu Wu, Jing Lei, Lu Hou 等CVPR 2026 · 被引用 4 次
- Latent Visual ReasoningBangzheng Li, Ximeng Sun, Jiang Liu, Ze Wang 等ICLR 2026 · 被引用 80 次
- GRIT: Teaching MLLMs to Think with ImagesYue Fan, Xuehai He, Diji Yang, Kaizhi Zheng 等NeurIPS 2025 · 被引用 132 次
- VisRef: Visual Refocusing while Thinking Improves Test-Time Scaling in Multi-Modal Large Reasoning ModelsSoumya Suvra Ghosal, Youngeun Kim, Zhuowei Li, Ritwick Chaudhry 等CVPR 2026
