FAFO: Lossy KV Cache Compression for Lossless Inference Acceleration via Draftless Fumble Decoding
Hoang Anh Duy Le, Shaochen (Henry) Zhong, Yifan Lu, Yingtong Dou, Jiayi Yuan, Yu-Neng Chuang, Xiran Fan, Guanchu Wang, Yuzhong Chen, Xia Hu
Abstract
Lossy KV cache compression is a well-explored subfield of machine learning efficiency, with improved latency being one of its major gains. However, lossy compression techniques can fumble from time to time, exhibiting various, and often catastrophic, failure patterns that are not only difficult to resolve but sometimes even hard to identify, making direct deployment of models with compressed KV cache a risky endeavor. In this work, we explore a way to preserve lossless generation quality while still benefiting from the acceleration provided by KV cache compression. Specifically, we draw inspiration from the n-gram candidate pool decoding paradigm where we purposely allow the model to Fumble Around with compressed KV cache to generate multiple lossy "n-gram guesses", while in parallel Find Out via lossless verification in the same forward pass. From a conceptual standpoint, our proposed framework is compatible with all typical static or dynamic KV cache compression methods from the token dropping realm, thus opening up a new avenue for the stagnant n-gram decoding paradigm. Practically, we show that this framework presents many useful traits that similar draftless baselines (e.g., Self-Speculative Decoding) cannot achieve, such as requiring only one set of KV cache and being far less sensitive to model, task, and input-length scenarios. Our comprehensive empirical results show FAFO provides 1.20-2.71× latency speedup over the original model, while consistently outperforming other lossless + draftless solutions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a3a06f09-dc87-4b1a-a0c4-d326846c57eeCited by top-tier papers1
Ask how each one uses itBuilds on24
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 1,472 citations
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen et al.NeurIPS 2023 · 1,003 citations
- Fast Model Editing at ScaleEric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn et al.ICLR 2022 · 527 citations
Related papers
- Block Verification Accelerates Speculative DecodingZiteng Sun, Uri Mendlovic, Yaniv Leviathan, Asaf Aharoni et al.ICLR 2025
- The Pitfalls of KV Cache CompressionAlex Chen, Renato Lui Geh, Aditya Grover, Guy Van den Broeck et al.ACL 2026 · 5 citations
- Value-Guided KV Compression for LLMs via Approximated CUR DecompositionAyan Sengupta, Siddhant Chaudhary, Tanmoy ChakrabortyNeurIPS 2025 · 6 citations
- Speculate Deep and Accurate: Lossless and Training-Free Acceleration for Offloaded LLMs via Substitute Speculative DecodingPei-Shuo Wang, Jian-Jia Chen, Chun-Che Yang, Chi-Chih Chang et al.NeurIPS 2025 · 1 citation
- Vegas: Self-Speculative Decoding with Verification-Guided Sparse AttentionYikang Yue, Yuqi Xue, Jian HuangICML 2026 · 2 citations
