Crystal: Introspective Reasoners Reinforced with Self-Feedback
Jiacheng Liu, Ramakanth Pasunuru, Hannaneh Hajishirzi, Yejin Choi, Asli Celikyilmaz
摘要
Extensive work has shown that the performance and interpretability of commonsense reasoning can be improved via knowledge-augmented reasoning methods, where the knowledge that underpins the reasoning process is explicitly verbalized and utilized. However, existing implementations, including "chain-of-thought" and its variants, fall short in capturing the introspective nature of knowledge required in commonsense reasoning, and in accounting for the mutual adaptation between the generation and utilization of knowledge. We propose a novel method to develop an introspective commonsense reasoner, CRYSTAL. To tackle commonsense problems, it first introspects for knowledge statements related to the given question, and subsequently makes an informed prediction that is grounded in the previously introspected knowledge. The knowledge introspection and knowledge-grounded reasoning modes of the model are tuned via reinforcement learning to mutually adapt, where the reward derives from the feedback given by the model itself. Experiments show that CRYSTAL significantly outperforms both the standard supervised finetuning and chain-of-thought distilled methods, and enhances the transparency of the commonsense reasoning process. Our work ultimately validates the feasibility and potential of reinforcing a neural model with self-feedback. 1 * Work done as a visiting researcher at FAIR, Meta.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference FeedbackHamish Ivison, Yizhong Wang, Jiacheng Liu, Zeqiu Wu 等NeurIPS 2024 · 被引用 124 次
- Grokking of Implicit Reasoning in Transformers: A Mechanistic Journey to the Edge of GeneralizationBoshi Wang, Xiang Yue, Yu Su, Huan SunNeurIPS 2024 · 被引用 48 次
- Navigate through Enigmatic Labyrinth A Survey of Chain of Thought Reasoning: Advances, Frontiers and FutureZheng Chu, Jingchang Chen, Qianglong Chen, Weijiang Yu 等ACL 2024 · 被引用 36 次
- SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated ResponsesDongwei Jiang, Jingyu Zhang, Orion Weller, Nathaniel Weir 等AAAI 2025 · 被引用 8 次
- Compositional Generalization from Learned Skills via CoT Training: A Theoretical and Structural Analysis for ReasoningXinhao Yao, Ruifeng Ren, Yun Liao, Lizhong Ding 等ICLR 2026 · 被引用 6 次
它引用的顶会 Paper19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
相关 Paper
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 被引用 1,126 次
- Why and How LLMs Benefit from Knowledge Introspection in Commonsense ReasoningChengfeng Zhao, Shizhu He, Shanshan Jiang, Bin Dong 等EMNLP 2025
- Thinking Out Loud: Do Reasoning Models Know When They're Right?Qingcheng Zeng, Weihao Xuan, Leyang Cui, Rob VoigtEMNLP 2025 · 被引用 1 次
- Beyond Accuracy: Dissecting Mathematical Reasoning for LLMs Under Reinforcement LearningJiayu Wang, Yifei Ming, Zixuan Ke, Caiming Xiong 等NeurIPS 2025 · 被引用 7 次
- Rainier: Reinforced Knowledge Introspector for Commonsense Question AnsweringJiacheng Liu, Skyler Hallinan, Ximing Lu, Pengfei He 等EMNLP 2022 · 被引用 31 次
