Head-Aware KV Cache Compression for Efficient Visual Autoregressive Modeling
Ziran Qin, Youru Lv, Mingbao Lin, Hang Guo, Zeren Zhang, Danping Zou, Weiyao Lin
摘要
Visual Autoregressive (VAR) models adopt a next-scale prediction paradigm, offering high-quality content generation with substantially fewer decoding steps. However, existing VAR models suffer from significant attention complexity and severe memory overhead due to the accumulation of keyvalue (KV) caches across scales. In this paper, we tackle this challenge by introducing KV cache compression into the next-scale generation paradigm. We begin with a crucial observation: attention heads in VAR models can be divided into two functionally distinct categories: Contextual Heads focus on maintaining semantic consistency, while Structural Heads are responsible for preserving spatial coherence. This structural divergence causes existing one-size-fits-all compression methods to perform poorly on VAR models. To address this, we propose HACK, a training-free Head-Aware KV cache Compression frameworK. HACK utilizes an offline classification scheme to separate head types, enabling it to apply pattern-specific compression strategies with asymmetric cache budgets for each category. By doing so, HACK effectively constrains the average KV cache length within a fixed budget B, reducing the theoretical attention complexity from O(n 4 ) to O(Bn 2 ). Extensive experiments on multiple VAR models across text-to-image and class-conditional tasks validate the effectiveness and generalizability of HACK. It achieves up to 70% KV cache compression without degrading output quality, resulting in memory savings and faster inference. For example, HACK provides a 1.75× memory reduction and a 1.57× speedup on Infinity-8B.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache CompressionKunjun Li, Zigeng Chen, Cheng-Yen Yang, Jenq-Neng HwangNeurIPS 2025 · 被引用 23 次
- Progressive Supernet Training for Efficient Visual Autoregressive ModelingXiaoyue Chen, Yuling Shi, Kaiyuan Li, Huandong Wang 等CVPR 2026 · 被引用 7 次
- ToProVAR: Efficient Visual Autoregressive Modeling via Tri-Dimensional Entropy-Aware Semantic Analysis and Sparsity OptimizationJiayu Chen, Ruoyu Lin, Zihao Zheng, Jingxin Li 等ICLR 2026 · 被引用 5 次
- SparVAR: Exploring Sparsity in Visual AutoRegressive Modeling for Training-Free AccelerationZekun Li, Ning Wang, Tongxin Bai, Changwang Mei 等CVPR 2026 · 被引用 4 次
- RADAR: VQ-VAE Decoder of VAR is a Good Student for Restoring Against Degradation by AccelerationZiyang Wang, Yue Zhang, Mingdao Wang, Yasen Zhang 等CVPR 2026
它引用的顶会 Paper22
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
相关 Paper
- AMS-KV: Adaptive KV Caching in Multi-Scale Visual Autoregressive TransformersBoxun Xu, Yu Wang, Zihu Wang, Peng LiAAAI 2026 · 被引用 2 次
- Entropy-Aware Dynamic KV Cache Sparsification for Autoregressive Image Generation and EditingTong Tong, LING XING, Linjie Li, Rui Yan 等ICML 2026
- HybridKV: Hybrid KV Cache Compression for Efficient Multimodal Large Language Model InferenceBowen Zeng, Feiyang Ren, Jun Zhang, Xiaoling Gu 等ACL 2026 · 被引用 5 次
- FasterVAR: Plug-and-Play Acceleration for Visual Autoregressive ModelsSenmao Li, Kai Wang, Salman Khan, Fahad Khan 等ICML 2026 · 被引用 2 次
- HeteroCache: A Dynamic Retrieval Approach to Heterogeneous KV Cache Compression for Long-Context LLM InferenceZhiyuan Shi, Qibo Qiu, Feng Xue, Zhonglin Jiang 等ACL 2026 · 被引用 1 次
