Critique-Guided Distillation for Robust Reasoning via Refinement
Berkcan Kapusuzoglu, Supriyo Chakraborty, Zain Sarwar, Michael Lee, Sambit Sahu
摘要
Supervised fine-tuning with expert demonstrations often produces models that imitate outputs without internalizing the reasoning processes needed for robust generalization. While critiquebased approaches show promise, training models to generate critiques directly, such as Critique Fine-Tuning (CFT), can lead to output-format drift and degradation of general capabilities. We propose CRITIQUE-GUIDED DISTILLATION (CGD), a training framework that decouples critique consumption from critique generation. During fine-tuning, the student is trained to refine flawed responses conditioned on teacher critiques. CGD treats critiques as a training-time-only supervision signal, encouraging internalization of erroraware reasoning: critiques guide learning but are absent at inference. Controlled ablations confirm that these reasoning gains are directly driven by the specificity and relevance of the teacher's feedback. Across five model families, CGD consistently outperforms CFT and standard distillation on mathematical reasoning benchmarks, yielding 7% average improvements and gains of up to +15.0% on AMC23 and +12.2% on MATH-500. On challenging competition problems such as AIME24 and AIME25, CGD achieves substantially higher Pass@1 and stronger performance at low Pass@k, indicating improved reasoning quality per sample. Importantly, CGD preserves general instruction-following capabilities where CFT degrades significantly (-21.3% on IFEval). These results position CGD as a practical and compute-efficient intermediate training paradigm for reasoning-centric tasks without introducing architectural inference-time overhead.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper21
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer 等NeurIPS 2022 · 被引用 2,039 次
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach 等ICLR 2022 · 被引用 1,976 次
相关 Paper
- Critique-Coder: Enhancing Coder Models by Critique Reinforcement LearningChi Ruan, Dongfu Jiang, Yubo Wang, Wenhu ChenICLR 2026 · 被引用 7 次
- RLKD: Distilling LLMs' Reasoning via Reinforcement LearningShicheng Xu, Liang Pang, Yunchang Zhu, Jia Gu 等AAAI 2026 · 被引用 2 次
- AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement LearningYang Chen, Zhuolin Yang, Zihan Liu, Chankyu Lee 等NeurIPS 2025 · 被引用 79 次
- Explain in Your Own Words: Improving Reasoning via Token-Selective Dual Knowledge DistillationMinsang Kim, Seung Jun BaekICLR 2026 · 被引用 15 次
- The Quality-Utility Paradox: Why High-Reward Data Impairs Small Model Mathematical ReasoningHaolong Qian, Xianliang Yang, Ma yinuo, Lirong Che 等ICML 2026
