HierKV: A Coarse-to-Fine Approach with Vision-Aware Banzhaf Values for Multi-Modal KV Cache Compression
Zeyi Lu, Jinpeng Wang, Yan Feng, Bin Chen, Jiawei Li, Yaowei Wang, Shu-Tao Xia
Abstract
Multimodal large language models (MLLMs) are increasingly deployed in Web mining applications such as search, retrieval-augmented assistants, and community question answering, where long and visually rich inputs substantially increase the memory and latency cost of attention through KV caching. Existing KV cache compression methods typically optimize a single granularity (e.g., token eviction or layer-wise allocation), which can misallocate cache capacity in MLLMs because layer roles, head functionalities, and instance-level visual reliance are highly heterogeneous. We propose HierKV, a hierarchical KV cache compression framework for MLLMs that performs coarse-to-fine budget allocation over layers, heads, and tokens. At the layer level, HierKV combines model-wise block importance estimated via Banzhaf values with an instance-wise visual sensitivity signal from attention perturbation to assign layer budgets. At the head and token levels, it prioritizes cross-modal interaction heads and performs budgeted top-K token retention within each head. Extensive experiments across multiple MLLMs and multimodal benchmarks show that HierKV achieves a favorable accuracy-latency-memory trade-off under strict cache budgets, delivering substantial decoding speedups and GPU memory reductions compared to strong baselines.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 3a48ffe2-4cec-44d8-8d47-3a967adcbb39Related papers
- HybridKV: Hybrid KV Cache Compression for Efficient Multimodal Large Language Model InferenceBowen Zeng, Feiyang Ren, Jun Zhang, Xiaoling Gu et al.ACL 2026 · 5 citations
- Revisiting Multimodal KV Cache Compression: A Frequency-Domain-Guided Outlier-KV-Aware ApproachYaoxin Yang, Peng Ye, Xudong Tan, Chongjun Tu et al.CVPR 2026 · 5 citations
- MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context InferenceKunxi Li, Zhonghua Jiang, Zhouzhou Shen, Zhaode Wang et al.ACL 2025
- EAKV: An Entropy-Driven Adaptive KV Compression Framework for Long Video UnderstandingHengrui Hu, Jingyu Li, Juntao Liang, Guanyu Chen et al.ICML 2026
- Mixing Importance with Diversity: Joint Optimization for KV Cache Compression in Large Vision-Language ModelsXuyang Liu, Xiyan Gui, Yuchao Zhang, Linfeng ZhangICLR 2026 · 16 citations
