CoIn: Coverage and Informativeness-Guided Token Reduction for Efficient Large Multimodal Models
Chenxi Du, Yongheng Deng, Jiani Liu, Yujia Zhang, Xi Chen, Ju Ren
摘要
Large Multimodal Models (LMMs) have shown remarkable success in visual understanding tasks. LMMs encode visual and textual inputs into tokens, which are then processed by Large Language Models (LLMs). However, the large number of visual tokens poses a major bottleneck for inference efficiency and memory usage. Reducing visual tokens is a promising training-free solution, but existing methods remain limited. Importance-based approaches suffer from poor generalization, are incompatible with kernel-level inference optimizations, and only consider information from a single modality. Diversity-based strategies typically focus on pairwise token redundancy and treat all tokens as equally important. Recent attempts to sequentially combine importance and diversity criteria still fail to address the intrinsic drawbacks of their underlying metrics. To address these limitations, we reformulate visual token reduction as an optimal subset selection problem jointly guided by two complementary objectives: informativeness and coverage. Informativeness is quantified through per-token intrinsic saliency and visual–textual alignment, while coverage is enforced via a volume-based subset selection criterion that ensures global representativeness in the visual feature space.This joint formulation effectively integrates visual saliency, cross-modal alignment, and global coverage in an end-to-end token selection process, yielding a computationally efficient, model-agnostic framework compatible with modern inference accelerators. Extensive experiments demonstrate that CoIn substantially reduces computation and memory cost while maintaining strong task performance. We will release our code once accepted.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu 等NeurIPS 2022 · 被引用 2,727 次
- Evaluating Object Hallucination in Large Vision-Language ModelsYifan Li, Yifan Du, Kun Zhou, Jinpeng Wang 等EMNLP 2023 · 被引用 344 次
相关 Paper
- SPLIT-VLM: Salience-Guided Partitioning towards Local Coverage for Importance-Aware Token Dropping in Vision-Language ModelsSeungil Lee, Gilha lee, Hyun KimICML 2026
- Task-Related Token Compression in Multimodal Large Language Models from an Explainability PerspectiveLei Lei, Jie Gu, Xiaokang Ma, Chu Tang 等ICLR 2026 · 被引用 3 次
- MMTok: Multimodal Coverage Maximization for Efficient Inference of VLMsSixun Dong, Juhua Hu, Mian Zhang, Ming Yin 等ICLR 2026 · 被引用 33 次
- VisionTrim: Unified Vision Token Compression for Training-Free MLLM AccelerationHanxun Yu, Wentong Li, Xuan Qu, Song Wang 等ICLR 2026 · 被引用 17 次
- AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and PruningYiwu Zhong, Zhuoming Liu, Yin Li, Liwei WangICCV 2025 · 被引用 1 次
