LaCache: Ladder-Shaped KV Caching for Efficient Long-Context Modeling of Large Language Models
Dachuan Shi, Yonggan Fu, Xiangchi Yuan, Zhongzhi Yu, Haoran You, Sixu Li, Xin Dong, Jan Kautz, Pavlo Molchanov, Yingyan Celine Lin
Abstract
Recent advancements in Large Language Models (LLMs) have spurred interest in numerous applications requiring robust long-range capabilities, essential for processing extensive input contexts and continuously generating extended outputs. As sequence lengths increase, the number of Key-Value (KV) pairs in LLMs escalates, creating a significant efficiency bottleneck. In this paper, we propose a new KV cache optimization paradigm called LaCache, a training-free method for efficient and accurate generative inference of LLMs. LaCache enables LLMs to simultaneously address both of the critical challenges in longrange modeling: robust long-range capabilities and continuous generation without running outof-memory (OOM). Specifically, LaCache integrates two key innovations: (1) a ladder-shaped KV cache pattern that stores KV pairs not only sequentially (left-to-right within each layer) but also across layers (from shallow to deep), providing an extended span for capturing long-range dependencies under a fixed storage budget, thereby boosting long-range capabilities; and (2) an iterative compaction mechanism that progressively compresses older caches, freeing up space for new tokens within a fixed cache size. This token distance-based dynamic compression enables more effective continuous generation under constrained cache budgets. Experiments across various tasks, benchmarks, and LLM models consistently validate LaCache's effectiveness in enhancing LLMs' long-range capabilities. Our code is available at https://github.com/GATECH-EIC/LaCache .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3d69a719-8ac0-42e5-a5c0-df370a74fb9fCited by top-tier papers7
- SwiReasoning: Switch-Thinking in Latent and Explicit for Pareto-Superior Reasoning LLMsDachuan Shi, Abedelkadir Asi, Keying Li, Xiangchi Yuan et al.ICLR 2026 · 22 citations
- ThinKV: Thought-Adaptive KV Cache Compression for Efficient Reasoning ModelsAkshat Ramachandran, Marina Neseem, Charbel Sakr, Rangharajan Venkatesan et al.ICLR 2026 · 19 citations
- Superficial Self-Improved Reasoners Benefit from Model MergingXiangchi Yuan, Chunhui Zhang, Zheyuan Liu, Dachuan Shi et al.EMNLP 2025 · 15 citations
- ReST-KV: Robust KV Cache Eviction with Layer-wise Output Reconstruction and Spatial-Temporal SmoothingYongqi An, Chang Lu, Kuan Zhu, Tao Yu et al.ICLR 2026 · 11 citations
- FreqKV: Key-Value Compression in Frequency Domain for Context Window ExtensionJushi Kai, Yixuan Wang, Boyi Zeng, Haoli Bai et al.ICLR 2026 · 7 citations
Builds on18
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- SnapKV: LLM Knows What You are Looking for Before GenerationYuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh et al.NeurIPS 2024 · 1,019 citations
Related papers
- FreeKV: Boosting KV Cache Retrieval for Efficient LLM InferenceGuangda Liu, Chengwei Li, Zhenyu Ning, Jing Lin et al.ICLR 2026 · 17 citations
- RefreshKV: Updating Small KV Cache During Long-form GenerationFangyuan Xu, Tanya Goyal, Eunsol ChoiACL 2025 · 6 citations
- Dynamic Memory Compression: Retrofitting LLMs for Accelerated InferencePiotr Nawrot, Adrian Lancucki, Marcin Chochowski, David Tarjan et al.ICML 2024 · 106 citations
- Homogeneous Keys, Heterogeneous Values: Exploiting Local KV Cache Asymmetry for Long-Context LLMsWanyun Cui, Mingwei XuNeurIPS 2025 · 7 citations
- MiniCache: KV Cache Compression in Depth Dimension for Large Language ModelsAkide Liu, Jing Liu, Zizheng Pan, Yefei He et al.NeurIPS 2024 · 160 citations
