Lune

KDD2026Top-tier venue

HierKV: A Coarse-to-Fine Approach with Vision-Aware Banzhaf Values for Multi-Modal KV Cache Compression

Zeyi Lu, Jinpeng Wang, Yan Feng, Bin Chen, Jiawei Li, Yaowei Wang, Shu-Tao Xia

2026Year

Abstract

Multimodal large language models (MLLMs) are increasingly deployed in Web mining applications such as search, retrieval-augmented assistants, and community question answering, where long and visually rich inputs substantially increase the memory and latency cost of attention through KV caching. Existing KV cache compression methods typically optimize a single granularity (e.g., token eviction or layer-wise allocation), which can misallocate cache capacity in MLLMs because layer roles, head functionalities, and instance-level visual reliance are highly heterogeneous. We propose HierKV, a hierarchical KV cache compression framework for MLLMs that performs coarse-to-fine budget allocation over layers, heads, and tokens. At the layer level, HierKV combines model-wise block importance estimated via Banzhaf values with an instance-wise visual sensitivity signal from attention perturbation to assign layer budgets. At the head and token levels, it prioritizes cross-modal interaction heads and performs budgeted top-K token retention within each head. Extensive experiments across multiple MLLMs and multimodal benchmarks show that HierKV achieves a favorable accuracy-latency-memory trade-off under strict cache budgets, delivering substantial decoding speedups and GPU memory reductions compared to strong baselines.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get 3a48ffe2-4cec-44d8-8d47-3a967adcbb39

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines