CATP: Contextually Adaptive Token Pruning for Efficient and Enhanced Multimodal In-Context Learning
Yanshu Li, Jianjiang Yang, Zhennan Shen, Ligong Han, Haoyan Xu, Ruixiang Tang
Abstract
Modern large vision-language models (LVLMs) convert each input image into a large set of tokens that far outnumber the text tokens. Although this improves visual perception, it also introduces severe image token redundancy. Because image tokens contain sparse information, many contribute little to reasoning but greatly increase inference cost. Recent image token pruning methods address this issue by identifying important tokens and removing the rest. These methods improve efficiency with only small performance drops. However, most of them focus on single-image tasks and overlook multimodal in-context learning (ICL), where redundancy is higher and efficiency is more important. Redundant tokens weaken the advantage of multimodal ICL for rapid domain adaptation and lead to unstable performance. When existing pruning methods are applied in this setting, they cause large accuracy drops, which exposes a clear gap and the need for new approaches. To address this, we propose Contextually Adaptive Token Pruning (CATP), a training-free pruning method designed for multimodal ICL. CATP uses two stages of progressive pruning that fully reflect the complex cross-modal interactions in the input sequence. After removing 77.8% of the image tokens, CATP achieves an average performance gain of 0.6% over the vanilla model on four LVLMs and eight benchmarks, clearly outperforming all baselines. At the same time, it improves efficiency by reducing inference latency by an average of 10.78%. CATP strengthens the practical value of multimodal ICL and lays the foundation for future progress in interleaved image-text settings. * Corresponding Author. Q: What is placed on the plate in front of the little girl in dark blue? A: Nothing. Q:What should the driver do when driving in front of this sign? A: Stop the car. Q : What object is the animal playing with? A: Ice block. Q : What condiment is placed on the table and next to the birthday cake? A: 3-shot in-context demonstrations (ICDs) Query sample (a) (b) Attention-based method (c) Diversity-based method (d) Contextually Adaptive Token Pruning (Ours)
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 71e1d2b0-9119-43b2-af21-aed4d7778437Cited by top-tier papers4
- ChartVerse: Scaling Chart Reasoning via Reliable Programmatic Synthesis from ScratchZheng Liu, Honglin Lin, Xiaoyang Wang, Xin Gao et al.ACL 2026 · 6 citations
- DEALT: LLM-driven Diversity-Enhanced Data Augmentation for Long-Tail Text ClassificationWayne Lu, Xiaoxi CuiAAAI 2026 · 2 citations
- VFLowOpt: A Token Pruning Framework for LMMs with Visual Information Flow-Guided OptimizationSihan Yang, Runsen Xu, Chenhang Cui, Tai Wang et al.ICCV 2025 · 1 citation
- From Blind Transfer to Wise Selection: Prototype-Driven Neighbor-Domain Adaptation for Fake News DetectionWayne Lu, Yiheng LiAAAI 2026
Builds on19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid InferenceZhihang Lin, Mingbao Lin, Luxi Lin, Rongrong JiAAAI 2025 · 121 citations
- Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMsQizhe Zhang, Mengzhen Liu, Lichen Li, Ming Lu et al.NeurIPS 2025 · 104 citations
Related papers
- Instruction-Guided Cross-Modal Clustering for Training-Free Visual Token Pruning in Vision-Language ModelsYunqian Yu, Biao Chen, Yunya Zhang, Tonglan Xie et al.AAAI 2026
- LVLM_CSP: Accelerating Large Vision Language Models via Clustering, Scattering, and Pruning for Reasoning SegmentationHanning Chen, Yang Ni, Wenjun Huang, Hyunwoo Oh et al.ACM MM 2025 · 1 citation
- AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and PruningYiwu Zhong, Zhuoming Liu, Yin Li, Liwei WangICCV 2025 · 1 citation
- QuietPrune: Query-Guided Early Token Pruning for Vision-Language ModelsTianxiao Gao, Shanwei Zhao, Shuo Fang, Shiai Zhu et al.CVPR 2026
- TOP-RL: Task-Optimized Progressive Token Pruning with Reinforcement Learning for Vision Language ModelsHengyi Wang, Weiying Xie, Hui Jiang, Yaotao Wei et al.AAAI 2026
