When Efficiency Meets Safety: A Benchmark Security Analysis of KV Cache Compression in Large Language Models
Xiaoxiao Ma, Kuofeng Gao, Zeyi Lu, Wenxi Jiang, Hao Fang, Hao Wu, Bin Chen, Shu-Tao Xia
Abstract
Key-Value (KV) caching is widely used in large language models (LLMs) to enable longcontext inference efficiently, yet its security implications remain underexplored. We present the first systematic study of how KV cache compression interacts with jailbreak attacks, evaluating four model families under diverse jailbreak attacks. We identify a double-edged effect: (i) on one hand, compression can induce Accidental Robustness, where optimizationbased and encoding-based attacks fail due to Malicious Semantic Eviction, where attacks' own attention redirection reduces the malicious query's cache importance, and Gradient Mismatch where discrete compression operations break jailbreak optimization. (ii) On the other hand, Vulnerability Paradox arises under merging-based compression for humandesigned Attacks, where aggressive merging in shallow layers triggers functional head collapse, amplifying attack success rates. To address this, we propose Safe-CAM, a history-aware, perhead feedback merging strategy that prevents safety degradation while maintaining efficiency. Experiments show Safe-CAM fully restores safety (0% ASR) and improves benign task performance with minimal overhead. Our study highlights that KV cache compression is not only an efficiency mechanism but also a safetycritical prerequisite for deploying LLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 241116f5-e6e7-4664-a334-212ca031fe3dBuilds on20
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka et al.NeurIPS 2024 · 1,166 citations
- SnapKV: LLM Knows What You are Looking for Before GenerationYuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh et al.NeurIPS 2024 · 1,019 citations
- KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache QuantizationColeman Hooper, Sehoon Kim, Hiva Mohammadzadeh, Michael W. Mahoney et al.NeurIPS 2024 · 738 citations
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 722 citations
Related papers
- KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM InferenceYuxuan Tian, Zihan Wang, Yebo Peng, Aomufei Yuan et al.AAAI 2026
- RobustKV: Defending Large Language Models against Jailbreak Attacks via KV EvictionTanqiu Jiang, Zian Wang, Jiacheng Liang, Changjiang Li et al.ICLR 2025
- SafetyMem: Adaptive Jailbreak Defense via Dual-Component Safety MemoryHao Wang, Ziyi Ni, Huacan Wang, Pin Lyu et al.ACL 2026
- Shadow in the Cache: Unveiling and Mitigating Privacy Risks of KV-cache in LLM InferenceZhifan Luo, Shuo Shao, Su Zhang, Lijing Zhou et al.NDSS 2026 · 32 citations
- CaM: Cache Merging for Memory-efficient LLMs InferenceYuxin Zhang, Yuxuan Du, Gen Luo, Yunshan Zhong et al.ICML 2024 · 66 citations
