Rethinking Chain-of-Thought from the Perspective of Self-Training
Zongqian Wu, Baoduo Xu, Ruochen Cui, Mengmeng Zhan, Xiaofeng Zhu, Lei Feng
摘要
Chain-of-thought (CoT) reasoning has emerged as an effective approach for activating latent capabilities in LLMs. Interestingly, we observe that both CoT reasoning and self-training share the core objective: iteratively leveraging model-generated information to progressively reduce prediction uncertainty. Building on this insight, we propose a novel CoT framework to improve reasoning performance. Our framework integrates two key components: (i) a task-specific prompt module that optimizes the initial reasoning process, and (ii) an adaptive reasoning iteration module that dynamically refines the reasoning process and addresses the limitations of previous CoT approaches, i.e., over-reasoning and high similarity between consecutive reasoning iterations. Extensive experiments show that the proposed method achieves significant advantages in both performance and computational efficiency. Our code is available at: https://github.com/zongqianwu/ST-COT .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Unveiling the Entropy Dynamics of Chain-of-Thought ReasoningTING XU, Xu He, Yupu Lu, Jiankai Sun 等ICML 2026 · 被引用 2 次
- PRISMA: Preference-Reinforced Self-Training Approach for Interpretable Emotionally Intelligent Negotiation DialoguesPrajwal Vijay Kajare, Priyanshu Priya, Bikash Santra, Asif EkbalACL 2026
- Can Reasoning Path still be Effective as Input? Bridging Post-Reasoning to Chain-of-Thought CompressionChengzhengxu Li, Xiaoming Liu, Zhaohan Zhang, Shengchao Liu 等ACL 2026
它引用的顶会 Paper11
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- FlexMatch: Boosting Semi-Supervised Learning with Curriculum Pseudo LabelingBowen Zhang, Yidong Wang, Wenxin Hou, Hao Wu 等NeurIPS 2021 · 被引用 1,389 次
- Confidence Regularized Self-TrainingYang Zou, Zhiding Yu, Xiaofeng Liu, B. V. K. Vijaya Kumar 等ICCV 2019 · 被引用 901 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
相关 Paper
- Revisiting Chain-of-Thought in Code Generation: Do Language Models Need to Learn Reasoning before Coding?Renbiao Liu, Anqi Li, Chaoding Yang, Hui Sun 等ICML 2025
- Large Language Models Can Self-ImproveJiaxin Huang, Shixiang Gu, Le Hou, Yuexin Wu 等EMNLP 2023 · 被引用 184 次
- Compositional Generalization from Learned Skills via CoT Training: A Theoretical and Structural Analysis for ReasoningXinhao Yao, Ruifeng Ren, Yun Liao, Lizhong Ding 等ICLR 2026 · 被引用 6 次
- Why Prompt Design Matters and Works: A Complexity Analysis of Prompt Search Space in LLMsXiang Zhang, Juntai Cao, Chenyu You, Dujian DingACL 2025 · 被引用 21 次
- Activation Control for Efficiently Eliciting Long Chain-of-thought Ability of Language ModelsZekai Zhao, Qi Liu, Kun Zhou, Zihan Liu 等NeurIPS 2025 · 被引用 10 次
