VisionPulse: Dynamic Visual Sparsity for Efficient Multimodal Reasoning
Hengbo Xu, Shengjie Jin, Yanbiao Ma, Zhiwu Lu
摘要
With the rapid advancement of large multimodal models (LMMs), inference-time overhead has become a key bottleneck for real-world deployment. Existing methods typically prune visual tokens at prefill, assuming the required visual evidence remains static during reasoning. However, we empirically show that visual evidence is strongly step-dependent: only a sparse subset of visual tokens is critical at each decoding step, and the critical set evolves across reasoning. Furthermore, we identify a coupled bottleneck where redundant visual context can steer the model toward query-irrelevant regions, lengthening the reasoning trace. Guided by these insights, we propose VisionPulse , a step-wise visual token pruning framework during reasoning. VisionPulse computes a lightweight visual attention mass to estimate the step-wise retention budget by exploiting its strong positive correlation with LMMs' effective visual token usage and retain only the most critical tokens under this budget. By enforcing visual sparsity during reasoning, VisionPulse filters redundant visual context while preserving relevant visual evidence, shortening reasoning traces naturally. Extensive experiments show that VisionPulse only retains 5% of visual tokens per step with reasoning traces shortened by 11.2%, while keeping accuracy almost unchanged.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- MM-Vet: Evaluating Large Multimodal Models for Integrated CapabilitiesWeihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang 等ICML 2024 · 被引用 1,191 次
- Are We on the Right Way for Evaluating Large Vision-Language Models?Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang 等NeurIPS 2024 · 被引用 1,029 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Balanced Meta-Softmax for Long-Tailed Visual RecognitionJiawei Ren, Cunjun Yu, Shunan Sheng, Xiao Ma 等NeurIPS 2020 · 被引用 861 次
- Not All Images are Worth 16x16 Words: Dynamic Transformers for Efficient Image RecognitionYulin Wang, Rui Huang, Shiji Song, Zeyi Huang 等NeurIPS 2021 · 被引用 283 次
相关 Paper
- Efficient Multimodal Spatial Reasoning via Dynamic and Asymmetric RoutingYixian Shen, Qi Bi, Zihan Wang, Zhiheng Yang 等ICLR 2026
- LVLM_CSP: Accelerating Large Vision Language Models via Clustering, Scattering, and Pruning for Reasoning SegmentationHanning Chen, Yang Ni, Wenjun Huang, Hyunwoo Oh 等ACM MM 2025 · 被引用 1 次
- VFLowOpt: A Token Pruning Framework for LMMs with Visual Information Flow-Guided OptimizationSihan Yang, Runsen Xu, Chenhang Cui, Tai Wang 等ICCV 2025 · 被引用 1 次
- SparseVILA: Decoupling Visual Sparsity for Efficient VLM InferenceSamir Khaki, Junxian Guo, Jiaming Tang, Shang Yang 等ICCV 2025 · 被引用 3 次
- PRIM:Cooperative Dynamic Token Compression for Efficient Large Multimodal ModelsSong Li, yongping xiongICML 2026
