HierKV: A Coarse-to-Fine Approach with Vision-Aware Banzhaf Values for Multi-Modal KV Cache Compression
Zeyi Lu, Jinpeng Wang, Yan Feng, Bin Chen, Jiawei Li, Yaowei Wang, Shu-Tao Xia
摘要
Multimodal large language models (MLLMs) are increasingly deployed in Web mining applications such as search, retrieval-augmented assistants, and community question answering, where long and visually rich inputs substantially increase the memory and latency cost of attention through KV caching. Existing KV cache compression methods typically optimize a single granularity (e.g., token eviction or layer-wise allocation), which can misallocate cache capacity in MLLMs because layer roles, head functionalities, and instance-level visual reliance are highly heterogeneous. We propose HierKV, a hierarchical KV cache compression framework for MLLMs that performs coarse-to-fine budget allocation over layers, heads, and tokens. At the layer level, HierKV combines model-wise block importance estimated via Banzhaf values with an instance-wise visual sensitivity signal from attention perturbation to assign layer budgets. At the head and token levels, it prioritizes cross-modal interaction heads and performs budgeted top-K token retention within each head. Extensive experiments across multiple MLLMs and multimodal benchmarks show that HierKV achieves a favorable accuracy-latency-memory trade-off under strict cache budgets, delivering substantial decoding speedups and GPU memory reductions compared to strong baselines.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- HybridKV: Hybrid KV Cache Compression for Efficient Multimodal Large Language Model InferenceBowen Zeng, Feiyang Ren, Jun Zhang, Xiaoling Gu 等ACL 2026 · 被引用 5 次
- Revisiting Multimodal KV Cache Compression: A Frequency-Domain-Guided Outlier-KV-Aware ApproachYaoxin Yang, Peng Ye, Xudong Tan, Chongjun Tu 等CVPR 2026 · 被引用 5 次
- MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context InferenceKunxi Li, Zhonghua Jiang, Zhouzhou Shen, Zhaode Wang 等ACL 2025
- EAKV: An Entropy-Driven Adaptive KV Compression Framework for Long Video UnderstandingHengrui Hu, Jingyu Li, Juntao Liang, Guanyu Chen 等ICML 2026
- Mixing Importance with Diversity: Joint Optimization for KV Cache Compression in Large Vision-Language ModelsXuyang Liu, Xiyan Gui, Yuchao Zhang, Linfeng ZhangICLR 2026 · 被引用 16 次
