DTP: Delta-Guided Two Stage Pruning for Mamba-based Multimodal Large Language Models
Seong-Yeol Park, Kwon-Min Jung, Xianghua Piao, Yeong Hyeon Gu
Abstract
Multimodal large language models built on the Mamba architecture offer efficiency advantages, yet remain hampered by redundant visual tokens that inflate inference cost, with the prefill stage accounting for the majority of total inference time. We introduce Delta-guided Two stage Pruning (DTP), a method that progressively reduces token redundancy through selective pruning at early layer and complete pruning at late layer. Unlike Transformer-oriented pruning methods, our approach derives token importance directly from Mamba’s internal parameters. The statistical distribution of these importance scores, combined with implicit attention patterns, then provides the basis for determining both the pruning layers and the tokens to be removed. Extensive evaluation across diverse benchmarks shows that DTP cuts computation by nearly 50%, maintains higher task performance than existing pruning methods, and further achieves over a 35% reduction in prefill latency. Beyond efficiency, our analysis reveals previously underexplored behaviors of visual tokens within Mamba layers, suggesting a principled perspective for designing future pruning techniques in Mamba-based Multimodal Large Language Models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 799e8d37-16b6-401b-a9bc-27a2423fe8a9Cited by top-tier papers1
Ask how each one uses itBuilds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang et al.ICML 2024 · 1,725 citations
- Evaluating Object Hallucination in Large Vision-Language ModelsYifan Li, Yifan Du, Kun Zhou, Jinpeng Wang et al.EMNLP 2023 · 344 citations
Related papers
- Stop Looking for "Important Tokens" in Multimodal Language Models: Duplication Matters MoreZichen Wen, Yifeng Gao, Shaobo Wang, Junyuan Zhang et al.EMNLP 2025 · 4 citations
- D²Pruner: Debiased Importance and Structural Diversity for MLLM Token PruningEvelyn Zhang, Fufu Yu, Aoqi Wu, Zichen Wen et al.AAAI 2026 · 1 citation
- ATP-LLaVA: Adaptive Token Pruning for Large Vision Language ModelsXubing Ye, Yukang Gan, Yixiao Ge, Xiao-Ping Zhang et al.CVPR 2025
- Hi-Lo Prune: Look at What You'll Lose before Pruning with Hierarchical Token SelectionZixun Sun, Yubo Dong, Hehe Fan, Yi YangCVPR 2026
- DCP: Dual-Cue Pruning for Efficient Large Vision-Language ModelsLei Jiang, Zixun Zhang, Yuting Zeng, Chunzhao Xie et al.EMNLP 2025 · 2 citations
