CATP: Contextually Adaptive Token Pruning for Efficient and Enhanced Multimodal In-Context Learning
Yanshu Li, Jianjiang Yang, Zhennan Shen, Ligong Han, Haoyan Xu, Ruixiang Tang
摘要
Modern large vision-language models (LVLMs) convert each input image into a large set of tokens that far outnumber the text tokens. Although this improves visual perception, it also introduces severe image token redundancy. Because image tokens contain sparse information, many contribute little to reasoning but greatly increase inference cost. Recent image token pruning methods address this issue by identifying important tokens and removing the rest. These methods improve efficiency with only small performance drops. However, most of them focus on single-image tasks and overlook multimodal in-context learning (ICL), where redundancy is higher and efficiency is more important. Redundant tokens weaken the advantage of multimodal ICL for rapid domain adaptation and lead to unstable performance. When existing pruning methods are applied in this setting, they cause large accuracy drops, which exposes a clear gap and the need for new approaches. To address this, we propose Contextually Adaptive Token Pruning (CATP), a training-free pruning method designed for multimodal ICL. CATP uses two stages of progressive pruning that fully reflect the complex cross-modal interactions in the input sequence. After removing 77.8% of the image tokens, CATP achieves an average performance gain of 0.6% over the vanilla model on four LVLMs and eight benchmarks, clearly outperforming all baselines. At the same time, it improves efficiency by reducing inference latency by an average of 10.78%. CATP strengthens the practical value of multimodal ICL and lays the foundation for future progress in interleaved image-text settings. * Corresponding Author. Q: What is placed on the plate in front of the little girl in dark blue? A: Nothing. Q:What should the driver do when driving in front of this sign? A: Stop the car. Q : What object is the animal playing with? A: Ice block. Q : What condiment is placed on the table and next to the birthday cake? A: 3-shot in-context demonstrations (ICDs) Query sample (a) (b) Attention-based method (c) Diversity-based method (d) Contextually Adaptive Token Pruning (Ours)
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- ChartVerse: Scaling Chart Reasoning via Reliable Programmatic Synthesis from ScratchZheng Liu, Honglin Lin, Xiaoyang Wang, Xin Gao 等ACL 2026 · 被引用 6 次
- DEALT: LLM-driven Diversity-Enhanced Data Augmentation for Long-Tail Text ClassificationWayne Lu, Xiaoxi CuiAAAI 2026 · 被引用 2 次
- VFLowOpt: A Token Pruning Framework for LMMs with Visual Information Flow-Guided OptimizationSihan Yang, Runsen Xu, Chenhang Cui, Tai Wang 等ICCV 2025 · 被引用 1 次
- From Blind Transfer to Wise Selection: Prototype-Driven Neighbor-Domain Adaptation for Fake News DetectionWayne Lu, Yiheng LiAAAI 2026
它引用的顶会 Paper19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid InferenceZhihang Lin, Mingbao Lin, Luxi Lin, Rongrong JiAAAI 2025 · 被引用 121 次
- Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMsQizhe Zhang, Mengzhen Liu, Lichen Li, Ming Lu 等NeurIPS 2025 · 被引用 104 次
相关 Paper
- Instruction-Guided Cross-Modal Clustering for Training-Free Visual Token Pruning in Vision-Language ModelsYunqian Yu, Biao Chen, Yunya Zhang, Tonglan Xie 等AAAI 2026
- LVLM_CSP: Accelerating Large Vision Language Models via Clustering, Scattering, and Pruning for Reasoning SegmentationHanning Chen, Yang Ni, Wenjun Huang, Hyunwoo Oh 等ACM MM 2025 · 被引用 1 次
- AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and PruningYiwu Zhong, Zhuoming Liu, Yin Li, Liwei WangICCV 2025 · 被引用 1 次
- QuietPrune: Query-Guided Early Token Pruning for Vision-Language ModelsTianxiao Gao, Shanwei Zhao, Shuo Fang, Shiai Zhu 等CVPR 2026
- TOP-RL: Task-Optimized Progressive Token Pruning with Reinforcement Learning for Vision Language ModelsHengyi Wang, Weiying Xie, Hui Jiang, Yaotao Wei 等AAAI 2026
