Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models
Weihao Ye, Qiong Wu, Wenhao Lin, Yiyi Zhou
摘要
Recent progress in Multimodal Large Language Models (MLLMs) often use large image tokens to compensate the visual shortcoming of MLLMs, which not only exhibits obvious redundancy but also greatly exacerbates the already high computation. Token pruning is an effective solution for speeding up MLLMs, but when and how to drop tokens still remains a challenge. In this paper, we propose a novel and training-free approach for the effective visual token pruning of MLLMs, termed FitPrune, which can quickly produce a complete pruning recipe for MLLMs according to a predefined budget. Specifically, FitPrune considers token pruning as a statistical problem of MLLM and its objective is to find out an optimal pruning scheme that can minimize the divergence of the attention distributions before and after pruning. In practice, FitPrune can be quickly accomplished based on the attention statistics from a small batch of inference data, avoiding the expensive trials of MLLMs. According to the pruning recipe, an MLLM can directly remove the redundant visual tokens of different examples during inference. To validate FitPrune, we apply it to a set of recent MLLMs, including LLaVA-1.5, LLaVA-HR and LLaVA-NEXT, and conduct extensive experiments on a set of benchmarks. The experimental results show that our FitPrune can not only reduce the computational complexity to a large extent, while retaining high performance, e.g., -54.9% FLOPs for LLaVA-NEXT with only 0.5% accuracy drop. Notably, the pruning recipe can be obtained in about 5 minutes. Our code is available at https://github.com/ywh187/FitPrune .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper59
- Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMsQizhe Zhang, Mengzhen Liu, Lichen Li, Ming Lu 等NeurIPS 2025 · 被引用 104 次
- MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot ManipulationRongyu Zhang, Menghang Dong, Yuan Zhang, Liang Heng 等AAAI 2026 · 被引用 56 次
- Don't Just Chase "Highlighted Tokens" in MLLMs: Revisiting Visual Holistic Context RetentionXin Zou, Di Lu, Yizhou Wang, Yibo Yan 等NeurIPS 2025 · 被引用 49 次
- OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language ModelsKeda Tao, Kele Shao, Bohan Yu, Weiqiang Wang 等CVPR 2026 · 被引用 32 次
- SpecPrune-VLA: Accelerating Vision-Language-Action Models via Action-Aware Self-Speculative PruningHanzhen Wang, Jiaming Xu, Yushun Xiang, Jiayi Pan 等ICML 2026 · 被引用 32 次
它引用的顶会 Paper26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
相关 Paper
- What Kind of Visual Tokens Do We Need? Training-Free Visual Token Pruning for Multi-Modal Large Language Models from the Perspective of GraphYutao Jiang, Qiong Wu, Wenhao Lin, Wei Yu 等AAAI 2025 · 被引用 27 次
- TransPrune: Token Transition Pruning for Efficient Large Vision-Language ModelAo Li, Yuxiang Duan, Jinghui Zhang, Congbo Ma 等CVPR 2026 · 被引用 3 次
- VisiPruner: Decoding Discontinuous Cross-Modal Dynamics for Efficient Multimodal LLMsYingqi Fan, Anhao Zhao, Jinlan Fu, Junlong Tong 等EMNLP 2025 · 被引用 11 次
- ST3: Accelerating Multimodal Large Language Model by Spatial-Temporal Visual Token TrimmingJiedong Zhuang, Lu Lu, Ming Dai, Rui Hu 等AAAI 2025 · 被引用 1 次
- Instruction-Guided Cross-Modal Clustering for Training-Free Visual Token Pruning in Vision-Language ModelsYunqian Yu, Biao Chen, Yunya Zhang, Tonglan Xie 等AAAI 2026
