Think Thrice Before You Act: Progressive Thought Refinement in Large Language Models
Chengyu Du, Jinyi Han, Yizhou Ying, Aili Chen, Qianyu He, Haokun Zhao, Haoran Guo, Sirui Xia, Jiaqing Liang, Zulong Chen, Liangyue Li, Yanghua Xiao
摘要
Recent advancements in large language models (LLMs) have demonstrated that progressive refinement, rather than providing a single answer, results in more accurate and thoughtful outputs. However, existing methods often rely heavily on supervision signals to evaluate previous responses, making it difficult to assess output quality in more open-ended scenarios effectively. Additionally, these methods are typically designed for specific tasks, which limits their generalization to new domains. To address these limitations, we propose Progressive Thought Refinement (PTR), a framework that enables LLMs to refine their responses progressively. PTR operates in two phases: (1) Thought data construction stage: We propose a weak and strong model collaborative selection strategy to build a high-quality progressive refinement dataset to ensure logical consistency from thought to answers, and the answers are gradually refined in each round. (2) Thought-Mask Fine-Tuning Phase: We design a training structure to mask the "thought" and adjust loss weights to encourage LLMs to refine prior thought, teaching them to implicitly understand "how to improve" rather than "what is correct." Experimental results show that PTR significantly enhances LLM performance across ten diverse tasks (avg. from 49.6% to 53.5%) without task-specific fine-tuning. Notably, in more open-ended tasks, LLMs also demonstrate substantial improvements in the quality of responses beyond mere accuracy, suggesting that PTR truly teaches LLMs to self-improve over time.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Selective Expert Guidance for Effective and Diverse Exploration in Reinforcement Learning of LLMsZishang Jiang, Jinyi Han, Tingyun Li, Xinyi Wang 等ICLR 2026 · 被引用 7 次
- ReCot: Reflective Self-Correction Training for Mitigating Confirmation Bias in Large Vision-Language ModelsMengxue Qu, Yibo Hu, Kunyang Han, Yunchao Wei 等ICCV 2025 · 被引用 3 次
- Data-Efficient Selection via Grammatical Complexity in Continual Pre-training of Domain-Specific LLMsYizhou Ying, Geng Zhang, Cui Danxin, Chengyu Du 等EMNLP 2025
它引用的顶会 Paper25
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
相关 Paper
- A Stitch in Time Saves Nine: Proactive Self-Refinement for Language ModelsJinyi Han, Xinyi Wang, Haiquan Zhao, tingyun li 等ICLR 2026 · 被引用 2 次
- MoT: Memory-of-Thought Enables ChatGPT to Self-ImproveXiaonan Li, Xipeng QiuEMNLP 2023 · 被引用 16 次
- ReActR: Reasoning through Error-Activated Reflection for LLM Post-TrainingLina SunACL 2026
- Enabling Lanuguage Models to Implicitly Learn Self-ImprovementZiqi Wang, Le Hou, Tianjian Lu, Yuexin Wu 等ICLR 2024 · 被引用 2 次
- Task-Level Thinking Steps Help Large Language Models for Challenging Classification TaskChunhui Du, Jidong Tian, Haoran Liao, Jindou Chen 等EMNLP 2023 · 被引用 2 次
