PuMer: Pruning and Merging Tokens for Efficient Vision Language Models
Qingqing Cao, Bhargavi Paranjape, Hannaneh Hajishirzi
Abstract
Large-scale vision language (VL) models use Transformers to perform cross-modal interactions between the input text and image. These cross-modal interactions are computationally expensive and memory-intensive due to the quadratic complexity of processing the input image and text. We present PuMer 1 : a token reduction framework that uses text-informed Pruning and modality-aware Merging strategies to progressively reduce the tokens of input image and text, improving model inference speed and reducing memory footprint. PuMer learns to keep salient image tokens related to the input text and merges similar textual and visual tokens by adding lightweight token reducer modules at several cross-modal layers in the VL model. Training PuMer is mostly the same as finetuning the original VL model but faster. Our evaluation for two vision language models on four downstream VL tasks shows PuMer increases inference throughput by up to 2x and reduces memory footprint by over 50% while incurring less than a 1% accuracy drop. 2
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 98383c09-da0d-4bab-ac7c-3fcbda748fe7Cited by top-tier papers29
- VLA-Cache: Efficient Vision-Language-Action Manipulation via Adaptive Token CachingSiyu Xu, Yunke Wang, Chenghao Xia, Dihao Zhu et al.NeurIPS 2025 · 95 citations
- FastVGGT: Fast Visual Geometry TransformerYou Shen, Zhipeng Zhang, Yansong Qu, Xiawu Zheng et al.ICLR 2026 · 73 citations
- VL-JEPA: Joint Embedding Predictive Architecture for Vision-languageDelong Chen, Mustafa Shukor, Théo Moutakanni, Willy Chung et al.ICLR 2026 · 60 citations
- Accelerating Transformers with Spectrum-Preserving Token MergingChau Tran, Duy M. H. Nguyen, Manh-Duy Nguyen, TrungTin Nguyen et al.NeurIPS 2024 · 51 citations
- MobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced TrainingPavan Kumar Anasosalu Vasu, Hadi Pouransari, Fartash Faghri, Raviteja Vemulapalli et al.CVPR 2024 · 29 citations
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
- DynamicViT: Efficient Vision Transformers with Dynamic Token SparsificationYongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu et al.NeurIPS 2021 · 1,343 citations
- OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning FrameworkPeng Wang, An Yang, Rui Men, Junyang Lin et al.ICML 2022 · 1,058 citations
Related papers
- LLaVA-Prumerge: Adaptive Token Reduction for Efficient Large Multimodal ModelsYuzhang Shang, Mu Cai, Bingxin Xu, Yong Jae Lee et al.ICCV 2025 · 37 citations
- FOLDER: Accelerating Multi-Modal Large Language Models with Enhanced PerformanceHaicheng Wang, Zhemeng Yu, Gabriele Spadaro, Chen Ju et al.ICCV 2025 · 3 citations
- Skip-It? Theoretical Conditions for Layer Skipping in Vision–Language ModelsMax Hartman, Vidhata Jayaraman, Moulik Choraria, Akhil Bhimaraju et al.ICML 2026 · 1 citation
- iLLaVA: An Image is Worth Fewer Than 1/3 Input Tokens in Large Multimodal ModelsLianyu Hu, Liqing Gao, Fanhua Shang, Liang Wan et al.ICLR 2026 · 9 citations
- DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and InferenceAditya Kumar Singh, Hitesh Kandala, Pratik Prabhanjan Brahma, Zicheng Liu et al.CVPR 2026
