KaVa: Latent Reasoning via Compressed KV-Cache Distillation
Anna Kuzina, Maciej Pióro, Babak Ehteshami Bejnordi
摘要
Large Language Models (LLMs) excel at multi-step reasoning problems with explicit chain-of-thought (CoT), but verbose traces incur significant computational costs and memory overhead, and often carry redundant, stylistic artifacts. Latent reasoning has emerged as an efficient alternative that internalizes the thought process, but it suffers from a critical lack of supervision, limiting its effectiveness on complex, natural-language reasoning traces. In this work we propose KaVa, the first framework that bridges this gap by distilling knowledge directly from a compressed KV-cache of the teacher into a latent-reasoning student via self-distillation, leveraging the representational flexibility of continuous latent tokens to align stepwise KV trajectories. We show that the abstract, unstructured knowledge within compressed KV-cache, which lacks direct token correspondence, can serve as a rich supervisory signal for a latent reasoning student. Empirically, the approach consistently outperforms strong latent baselines, exhibits markedly smaller degradation from equation-only to natural-language traces, and scales to larger backbones while preserving efficiency. These results establish compressed KV-cache distillation as a scalable supervision signal for latent reasoning, combining the accuracy of CoT-trained teachers with the efficiency and deployability of latent inference.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- PAL: Program-aided Language ModelsLuyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon 等ICML 2023 · 被引用 700 次
- MetaMath: Bootstrap Your Own Mathematical Questions for Large Language ModelsLonghui Yu, Weisen Jiang, Han Shi, Jincheng Yu 等ICLR 2024 · 被引用 637 次
- Think before you speak: Training Language Models With Pause TokensSachin Goyal, Ziwei Ji, Ankit Singh Rawat, Aditya Krishna Menon 等ICLR 2024 · 被引用 240 次
- Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM InferenceHarry Dong, Xinyu Yang, Zhenyu Zhang, Zhangyang Wang 等ICML 2024 · 被引用 84 次
相关 Paper
- CODI: Compressing Chain-of-Thought into Continuous Space via Self-DistillationZhenyi Shen, Hanqi Yan, Linhai Zhang, Zhanghao Hu 等EMNLP 2025
- Internalizing Explicit Reasoning into Latent Space for Dense RetrievalJiajie Jin, Yanzhao Zhang, Mingxin Li, Dingkun Long 等SIGIR 2026
- Token Assorted: Mixing Latent and Text Tokens for Improved Language Model ReasoningDiJia Su, Hanlin Zhu, Yingchen Xu, Jiantao Jiao 等ICML 2025
- Mentor-KD: Making Small Language Models Better Multi-step ReasonersHojae Lee, Junho Kim, SangKeun LeeEMNLP 2024
- System-1.5 Reasoning: Traversal in Language and Latent Spaces with Dynamic ShortcutsXiaoqiang Wang, Suyuchen Wang, Yun Zhu, Bang LiuNeurIPS 2025 · 被引用 26 次
