DCP: Dual-Cue Pruning for Efficient Large Vision-Language Models
Lei Jiang, Zixun Zhang, Yuting Zeng, Chunzhao Xie, Tongxuan Liu, Zhen Li, Lechao Cheng, Xiaohua Xu
摘要
Large Vision-Language Models (LVLMs) achieve remarkable performance in multimodal tasks but suffer from high computational costs due to the large number of visual tokens. Existing pruning methods either apply after visual tokens enter the LLM or perform pre-pruning based solely on visual attention. Both fail to balance efficiency and semantic alignment, as post-pruning incurs redundant computation, while visual-only pre-pruning overlooks multimodal relevance. To address this limitation, we propose Dual-Cue Pruning (DCP), a novel cross-modal pruning framework that jointly considers textual semantics and visual selfattention. DCP consists of a text-aware computation module, which employs a gradientweighted attention mechanism to enhance textvisual alignment, and an image-aware computation module, which utilizes deep-layer selfattention distributions to retain essential structural information. By integrating both cues, DCP adaptively selects the most informative visual tokens, achieving efficient inference acceleration while maintaining strong task performance. Experimental results show that DCP can retain only 25% of the visual tokens, with a minimal performance degradation of 0.063% on LLaVA-1.5-13B, demonstrating its effectiveness in balancing efficiency and accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- RTPrune: Reading-Twice Inspired Token Pruning for Efficient DeepSeek-OCR InferenceBen Wan, Yan Feng, Zihan Tang, Weizhe Huang 等ICML 2026 · 被引用 3 次
- Open-World 3D Scene Graph Generation for Retrieval-Augmented ReasoningFei Yu, Quan Deng, Shengeng Tang, Yuehua Li 等AAAI 2026 · 被引用 2 次
- Decoupled Training with Local Reinforcement Fine-Tuning in Federated LearningYuting Ma, Lechao Cheng, Xiaohua XuICML 2026
- Reducing Token Redundancy in LVLMs: A Systematic Review of Token Pruning MethodsHanzhang Yuan, Mengxuan Hu, Wenhao Zhang, Tianlong Wang 等ACL 2026
- How Do LLMs and VLMs Understand Viewpoint Rotation Without Vision? An Interpretability StudyZhen Yang, Ping Jian, Zhongbin Guo, Zuming Zhang 等ACL 2026
它引用的顶会 Paper15
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu 等NeurIPS 2022 · 被引用 2,727 次
- Evaluating Object Hallucination in Large Vision-Language ModelsYifan Li, Yifan Du, Kun Zhou, Jinpeng Wang 等EMNLP 2023 · 被引用 344 次
- MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGIXiang Yue, Yuansheng Ni, Tianyu Zheng, Kai Zhang 等CVPR 2024 · 被引用 213 次
- Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMsShengbang Tong, Zhuang Liu, Yuexiang Zhai, Yi Ma 等CVPR 2024 · 被引用 111 次
相关 Paper
- Instruction-Guided Cross-Modal Clustering for Training-Free Visual Token Pruning in Vision-Language ModelsYunqian Yu, Biao Chen, Yunya Zhang, Tonglan Xie 等AAAI 2026
- Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMsQizhe Zhang, Aosong Cheng, Ming Lu, Renrui Zhang 等ICCV 2025 · 被引用 8 次
- VisiPruner: Decoding Discontinuous Cross-Modal Dynamics for Efficient Multimodal LLMsYingqi Fan, Anhao Zhao, Jinlan Fu, Junlong Tong 等EMNLP 2025 · 被引用 11 次
- DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and InferenceAditya Kumar Singh, Hitesh Kandala, Pratik Prabhanjan Brahma, Zicheng Liu 等CVPR 2026
- Don't Just Chase "Highlighted Tokens" in MLLMs: Revisiting Visual Holistic Context RetentionXin Zou, Di Lu, Yizhou Wang, Yibo Yan 等NeurIPS 2025 · 被引用 49 次
