Head-Aware KV Cache Compression for Efficient Visual Autoregressive Modeling
Ziran Qin, Youru Lv, Mingbao Lin, Hang Guo, Zeren Zhang, Danping Zou, Weiyao Lin
Abstract
Visual Autoregressive (VAR) models adopt a next-scale prediction paradigm, offering high-quality content generation with substantially fewer decoding steps. However, existing VAR models suffer from significant attention complexity and severe memory overhead due to the accumulation of keyvalue (KV) caches across scales. In this paper, we tackle this challenge by introducing KV cache compression into the next-scale generation paradigm. We begin with a crucial observation: attention heads in VAR models can be divided into two functionally distinct categories: Contextual Heads focus on maintaining semantic consistency, while Structural Heads are responsible for preserving spatial coherence. This structural divergence causes existing one-size-fits-all compression methods to perform poorly on VAR models. To address this, we propose HACK, a training-free Head-Aware KV cache Compression frameworK. HACK utilizes an offline classification scheme to separate head types, enabling it to apply pattern-specific compression strategies with asymmetric cache budgets for each category. By doing so, HACK effectively constrains the average KV cache length within a fixed budget B, reducing the theoretical attention complexity from O(n 4 ) to O(Bn 2 ). Extensive experiments on multiple VAR models across text-to-image and class-conditional tasks validate the effectiveness and generalizability of HACK. It achieves up to 70% KV cache compression without degrading output quality, resulting in memory savings and faster inference. For example, HACK provides a 1.75× memory reduction and a 1.57× speedup on Infinity-8B.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7b496dd2-5a98-47ea-8ab2-cdfac6eb737bCited by top-tier papers7
- Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache CompressionKunjun Li, Zigeng Chen, Cheng-Yen Yang, Jenq-Neng HwangNeurIPS 2025 · 23 citations
- Progressive Supernet Training for Efficient Visual Autoregressive ModelingXiaoyue Chen, Yuling Shi, Kaiyuan Li, Huandong Wang et al.CVPR 2026 · 7 citations
- ToProVAR: Efficient Visual Autoregressive Modeling via Tri-Dimensional Entropy-Aware Semantic Analysis and Sparsity OptimizationJiayu Chen, Ruoyu Lin, Zihao Zheng, Jingxin Li et al.ICLR 2026 · 5 citations
- SparVAR: Exploring Sparsity in Visual AutoRegressive Modeling for Training-Free AccelerationZekun Li, Ning Wang, Tongxin Bai, Changwang Mei et al.CVPR 2026 · 4 citations
- RADAR: VQ-VAE Decoder of VAR is a Good Student for Restoring Against Degradation by AccelerationZiyang Wang, Yue Zhang, Mingdao Wang, Yasen Zhang et al.CVPR 2026
Builds on22
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
Related papers
- AMS-KV: Adaptive KV Caching in Multi-Scale Visual Autoregressive TransformersBoxun Xu, Yu Wang, Zihu Wang, Peng LiAAAI 2026 · 2 citations
- Entropy-Aware Dynamic KV Cache Sparsification for Autoregressive Image Generation and EditingTong Tong, LING XING, Linjie Li, Rui Yan et al.ICML 2026
- HybridKV: Hybrid KV Cache Compression for Efficient Multimodal Large Language Model InferenceBowen Zeng, Feiyang Ren, Jun Zhang, Xiaoling Gu et al.ACL 2026 · 5 citations
- FasterVAR: Plug-and-Play Acceleration for Visual Autoregressive ModelsSenmao Li, Kai Wang, Salman Khan, Fahad Khan et al.ICML 2026 · 2 citations
- HeteroCache: A Dynamic Retrieval Approach to Heterogeneous KV Cache Compression for Long-Context LLM InferenceZhiyuan Shi, Qibo Qiu, Feng Xue, Zhonglin Jiang et al.ACL 2026 · 1 citation
