Seeing What Matters: A Training-Free Self-Guided Framework for Multimodal Detail Perception and Reasoning
Mingjie Ma, yichao ma, Zhong Yang, Guohui Li
摘要
Multimodal large language models (MLLMs) have achieved remarkable success on diverse visual-language tasks. However, fixed-resolution models face challenges in perceiving fine-grained visual details, particularly due to distracted attention and blurry vision . To address these issues, we propose SLoFo , a training-free and self-guided inference framework that mimics the human " S can- Lo cate- Fo cus" process. SLoFo first adopts a dual-branch mechanism to identify critical image regions: the Semantic branch constructs a gradient-based semantic relevance map, and the Structure branch estimates visual token uniqueness offering complementary and robust evidence. By combining both branches, SLoFo perceives and explicitly crop critical regions. During inference, with additional cropped sub-image, SLoFo applies a progressive visual token pruning strategy to improve attention focus on key areas while reducing computational overhead. Experiments on detail-sensitive and general-purpose benchmarks show that SLoFo consistently improves accuracy (+4.79% on TextVQA, +2.62% on GQA) and robustness (+4.60% on POPE-MSCOCO adversarial) without training or external modules.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
相关 Paper
- FOCUS: Internal MLLM Representations for Efficient Fine-Grained Visual Question AnsweringLiangyu Zhong, Fabio Rosenthal, Joachim Sicking, Fabian Hüger 等NeurIPS 2025 · 被引用 23 次
- Hi-Lo Prune: Look at What You'll Lose before Pruning with Hierarchical Token SelectionZixun Sun, Yubo Dong, Hehe Fan, Yi YangCVPR 2026
- AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and PruningYiwu Zhong, Zhuoming Liu, Yin Li, Liwei WangICCV 2025 · 被引用 1 次
- Blink: Dynamic Visual Token Resolution for Enhanced Multimodal UnderstandingYuchen Feng, Zhenyu Zhang, Naibin Gu, Yilong Chen 等CVPR 2026 · 被引用 2 次
- A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMsWangbo Zhao, Yizeng Han, Jiasheng Tang, Zhikai Li 等CVPR 2025
