TokenSkip: Controllable Chain-of-Thought Compression in LLMs
Heming Xia, Chak Tou Leong, Wenjie Wang, Yongqi Li, Wenjie Li
Abstract
Chain-of-Thought (CoT) has been proven effective in enhancing the reasoning capabilities of large language models (LLMs). Recent advancements, such as OpenAI's o1 and DeepSeek-R1, suggest that scaling up the length of CoT sequences during inference could further boost LLM reasoning performance. However, due to the autoregressive nature of LLM decoding, longer CoT outputs lead to a linear increase in inference latency, adversely affecting user experience, particularly when the CoT exceeds 10,000 tokens. To address this limitation, we analyze the semantic importance of tokens within CoT outputs and reveal that their contributions to reasoning vary. Building on this insight, we propose TokenSkip, a simple yet effective approach that enables LLMs to selectively skip less important tokens, allowing for controllable CoT compression. Extensive experiments across various models and tasks demonstrate the effectiveness of TokenSkip in reducing CoT token usage while preserving strong reasoning performance. Notably, when applied to Qwen2.5-14B-Instruct, TokenSkip reduces reasoning tokens by 40% (from 313 to 181) on GSM8K, with less than a 0.4% performance drop. We release our code and checkpoints in https: //github.com/hemingkx/TokenSkip .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers61
- Learn to Reason Efficiently with Adaptive Length-based Reward ShapingWei Liu, Ruochen Zhou, Yiyun Deng, Yuzhen Huang et al.ICLR 2026 · 88 citations
- InftyThink: Breaking the Length Limits of Long-Context Reasoning in Large Language ModelsYuchen Yan, Yongliang Shen, Yang Liu, Jin Jiang et al.ICLR 2026 · 48 citations
- Making Slow Thinking Faster: Compressing LLM Chain-of-Thought via Step EntropyZeju Li, Jianyuan Zhong, Ziyang Zheng, Xiangyu Wen et al.ICLR 2026 · 35 citations
- Ada-R1: Hybrid-CoT via Bi-Level Adaptive Reasoning OptimizationHaotian Luo, Haiying He, Yibo Wang, Jinluan Yang et al.NeurIPS 2025 · 29 citations
- Does Your Reasoning Model Implicitly Know When to Stop Thinking?Zixuan Huang, Xin Xia, Yuxi Ren, Jianbin Zheng et al.ICML 2026 · 21 citations
Builds on22
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
Related papers
- TokenSqueeze: Performance-Preserving Compression for Reasoning LLMsYuxiang Zhang, Zhengxu Yu, Weihang Pan, Zhongming Jin et al.NeurIPS 2025 · 6 citations
- Can Reasoning Path still be Effective as Input? Bridging Post-Reasoning to Chain-of-Thought CompressionChengzhengxu Li, Xiaoming Liu, Zhaohan Zhang, Shengchao Liu et al.ACL 2026
- Chain-of-Thought Compression Should Not Be Blind: V-Skip for Efficient Multimodal Reasoning via Dual-Path AnchoringDongxu Zhang, Yiding Sun, Cheng Tan, Wenbiao Yan et al.ACL 2026 · 18 citations
- ImgCoT: Compressing Long Chain of Thought into Compact Visual Tokens for Efficient Reasoning of Large Language ModelXiaoshu Chen, sihang zhou, KE LIANG, Taichun Zhou et al.ICML 2026 · 1 citation
- System-1.5 Reasoning: Traversal in Language and Latent Spaces with Dynamic ShortcutsXiaoqiang Wang, Suyuchen Wang, Yun Zhu, Bang LiuNeurIPS 2025 · 26 citations
