ConCISE: Confidence-guided Compression in Step-by-step Efficient Reasoning
Ziqing Qiao, Yongheng Deng, Jiali Zeng, Dong Wang, Lai Wei, Guanbo Wang, Fandong Meng, Jie Zhou, Ju Ren, Yaoxue Zhang
摘要
Large Reasoning Models (LRMs) perform strongly in complex reasoning tasks via Chainof-Thought (CoT) prompting, but often suffer from verbose outputs, increasing computational overhead. Existing fine-tuning-based compression methods either operate post-hoc pruning, risking disruption to reasoning coherence, or rely on sampling-based selection, which fails to remove redundant content thoroughly. To address these limitations, this work begins by framing two key patterns of redundant reflection in LRMs-Confidence Deficit, wherein the model reflects on correct intermediate steps, and Termination Delay, where reflection continues after a verified, confident answer-through a confidence-guided perspective. Based on this, we introduce CONCISE (Confidence-guided Compression In Step-bystep Efficient Reasoning), a framework designed to generate concise reasoning chains, integrating Confidence Injection to boost reasoning confidence, and Early Stopping to terminate reasoning when confidence is sufficient. Extensive experiments demonstrate that compared to baseline methods, fine-tuning LRMs on CONCISE-generated data yields a better balance between compression and task performance, reducing length by up to 50% under SimPO, while maintaining high task accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper25
- Dynamic Early Exit in Reasoning ModelsChenxu Yang, Qingyi Si, Yongjie Duan, Zheliang Zhu 等ICLR 2026 · 被引用 250 次
- S-GRPO: Early Exit via Reinforcement Learning in Reasoning ModelsMuzhi Dai, Chenxu Yang, Qingyi SiNeurIPS 2025 · 被引用 100 次
- Efficiently Scaling LLM Reasoning Programs with CertaindexYichao Fu, Junda Chen, Siqi Zhu, Zheyu Fu 等NeurIPS 2025 · 被引用 42 次
- Ada-R1: Hybrid-CoT via Bi-Level Adaptive Reasoning OptimizationHaotian Luo, Haiying He, Yibo Wang, Jinluan Yang 等NeurIPS 2025 · 被引用 29 次
- Does Your Reasoning Model Implicitly Know When to Stop Thinking?Zixuan Huang, Xin Xia, Yuxi Ren, Jianbin Zheng 等ICML 2026 · 被引用 21 次
它引用的顶会 Paper12
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- SimPO: Simple Preference Optimization with a Reference-Free RewardYu Meng, Mengzhou Xia, Danqi ChenNeurIPS 2024 · 被引用 1,203 次
- Training Language Models to Reason EfficientlyDaman Arora, Andrea ZanetteNeurIPS 2025 · 被引用 270 次
- Dynamic Early Exit in Reasoning ModelsChenxu Yang, Qingyi Si, Yongjie Duan, Zheliang Zhu 等ICLR 2026 · 被引用 250 次
相关 Paper
- ConPress: Learning Efficient Reasoning from Multi-Question Contextual PressureJie Deng, Shining Liang, Jun Li, Hongzhi Li 等ICML 2026 · 被引用 3 次
- Learning Graph Rationales to Compress Long Chains of Thought in Multimodal ReasoningYizhi Wang, Linan Yue, Deng-Bao Wang, Tong Wei 等KDD 2026
- Optimizing Length Compression in Large Reasoning ModelsZhengxiang Cheng, Dongping Chen, Mingyang Fu, Tianyi ZhouACL 2026 · 被引用 32 次
- Making Slow Thinking Faster: Compressing LLM Chain-of-Thought via Step EntropyZeju Li, Jianyuan Zhong, Ziyang Zheng, Xiangyu Wen 等ICLR 2026 · 被引用 35 次
- Your Reasoning Model Knows What Counts: Self-Guided Chain-of-Thought Pruning for Efficient ReasoningZi-Ao Ma, Xian-Ling Mao, Tian Lan, Chen Xu 等ACL 2026
