LightThinker: Thinking Step-by-Step Compression
Jintian Zhang, Yuqi Zhu, Mengshu Sun, Yujie Luo, Shuofei Qiao, Lun Du, Da Zheng, Huajun Chen, Ningyu Zhang
Abstract
Large language models (LLMs) have shown remarkable performance in complex reasoning tasks, but their efficiency is hindered by the substantial memory and computational costs associated with generating lengthy tokens. In this paper, we propose LightThinker, a novel method that enables LLMs to dynamically compress intermediate thoughts during reasoning. Inspired by human cognitive processes, Light-Thinker compresses verbose thought steps into compact representations and discards the original reasoning chains, thereby significantly reducing the number of tokens stored in the context window. This is achieved by training the model on when and how to perform compression through data construction, mapping hidden states to condensed gist tokens, and creating specialized attention masks. Additionally, we introduce the Dependency (Dep) metric to quantify the degree of compression by measuring the reliance on historical tokens during generation. Extensive experiments on four datasets and two models show that LightThinker reduces peak memory usage and inference time, while maintaining competitive accuracy. Our work provides a new direction for improving the efficiency of LLMs in complex reasoning tasks without sacrificing performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 049160a8-653e-44df-be42-34167b37a3f8Cited by top-tier papers15
- InftyThink: Breaking the Length Limits of Long-Context Reasoning in Large Language ModelsYuchen Yan, Yongliang Shen, Yang Liu, Jin Jiang et al.ICLR 2026 · 48 citations
- SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial ReasoningYang Liu, Ming Ma, Xiaomin Yu, Pengxiang Ding et al.NeurIPS 2025 · 46 citations
- NeuReasoner: Towards Explainable, Controllable, and Unified Reasoning via Mixture-of-NeuronsHaonan Dong, Kehan Jiang, Haoran Ye, Wenhao Zhu et al.ACL 2026 · 15 citations
- Render-of-Thought: Rendering Textual Chain-of-Thought as Images for Visual Latent ReasoningYifan Wang, Shiyu Li, Peiming Li, Xiaochen Yang et al.ACL 2026 · 14 citations
- LazyEviction: Lagged KV Eviction with Attention Pattern Observation for Efficient Long ReasoningHaoyue Zhang, Hualei Zhang, Xiaosong Ma, Jie Zhang et al.ACL 2026 · 7 citations
Builds on16
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- SnapKV: LLM Knows What You are Looking for Before GenerationYuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh et al.NeurIPS 2024 · 1,019 citations
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen et al.NeurIPS 2023 · 1,003 citations
- Learning to Compress Prompts with Gist TokensJesse Mu, Xiang Li, Noah D. GoodmanNeurIPS 2023 · 488 citations
Related papers
- Can Reasoning Path still be Effective as Input? Bridging Post-Reasoning to Chain-of-Thought CompressionChengzhengxu Li, Xiaoming Liu, Zhaohan Zhang, Shengchao Liu et al.ACL 2026
- DesireKV: Decoupling Sensitivity and Importance for Reasoning-Aware KV Cache CompressionPengyu Cheng, Jiacheng Wang, Tianle Chen, Bei Liu et al.AAAI 2026
- VeriThinker: Learning to Verify Makes Reasoning Model EfficientZigeng Chen, Xinyin Ma, Gongfan Fang, Ruonan Yu et al.NeurIPS 2025 · 30 citations
- Think Silently, Think Fast: Dynamic Latent Compression of LLM Reasoning ChainsWenhui Tan, Jiaze Li, Jianzhong Ju, Zhenbo Luo et al.NeurIPS 2025 · 103 citations
- ZoomR: Memory Efficient Reasoning through Multi-Granularity Key Value RetrievalDavid H. Yang, Yuxuan Zhu, Mohammad Mohammadi Amiri, Keerthiram Murugesan et al.ACL 2026
