QuietPrune: Query-Guided Early Token Pruning for Vision-Language Models
Tianxiao Gao, Shanwei Zhao, Shuo Fang, Shiai Zhu, Chenguang Ma
摘要
Vision-language models (VLMs) demonstrate powerful capabilities in multimodal tasks. However, the large number of visual tokens imposes a significant computational cost. In this paper, we propose QuietPrune, a QUery-guIded Early Token Pruning method to remove redundant visual tokens within VLMs, thereby enhancing computational efficiency. Unlike previous late pruning methods, we recognize that implementing early pruning within the vision transformer (ViT) can achieve benefits in both latency reduction and accuracy maintenance. To address the semantic loss problem in early pruning, we design a lightweight adapter by performing a inverse transformation of the projector in VLMs. The proposed adapter converts the contextual query into a visual domain [Q-CLS] (Query [CLS]) token, providing textual guidance for ViT pruning. During pruning, we further introduce a semi-structured pruning scheme based on visual-textual relevance. Specifically, we group spatially adjacent tokens to accommodate the visual token merging operation prevalent in mainstream VLMs. We use the mean attention scores between the [Q-CLS] token and the visual tokens as the relevance metric for each group, avoiding additional computation. Pruning is then applied at the group level based on the relevance score, preserving positional continuity. After pruning, we aggregate the redundant tokens into a single token to maintain context cues. Our method achieves up to 19.0% reduction in prefill latency while outperforming 4.2% in accuracy on the recent Qwen3-VL and InternVL3 series compared to existing late pruning methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- DynamicViT: Efficient Vision Transformers with Dynamic Token SparsificationYongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu 等NeurIPS 2021 · 被引用 1,343 次
相关 Paper
- Hi-Lo Prune: Look at What You'll Lose before Pruning with Hierarchical Token SelectionZixun Sun, Yubo Dong, Hehe Fan, Yi YangCVPR 2026
- ViTCoP: Accelerating Large Vision-Language Models via Visual and Textual Semantic Collaborative PruningWen Luo, Peng Chen, Xiaotao Huang, LiQun HuangAAAI 2026
- LearnPruner: Rethinking Attention-based Token Pruning in Vision Language ModelsRinyoichi Takezoe, Yaqian Li, Zi-Hao Bo, Anzhou Hou 等ICLR 2026 · 被引用 8 次
- HAWK: Head Importance-Aware Visual Token Pruning in Multimodal ModelsQihui Zhu, Tao Zhang, Yuchen Wang, Shuangwu Chen 等CVPR 2026 · 被引用 4 次
- TopV: Compatible Token Pruning with Inference Time Optimization for Fast and Low-Memory Multimodal Vision Language ModelCheng Yang, Yang Sui, Jinqi Xiao, Lingyi Huang 等CVPR 2025
