Generating Natural Language Proofs with Verifier-Guided Search
Kaiyu Yang, Jia Deng, Danqi Chen
摘要
Reasoning over natural language is a challenging problem in NLP. In this work, we focus on proof generation: Given a hypothesis and a set of supporting facts, the model generates a proof tree indicating how to derive the hypothesis from supporting facts. Compared to generating the entire proof in one shot, stepwise generation can better exploit the compositionality and generalize to longer proofs but has achieved limited success on real-world data. Existing stepwise methods struggle to generate proof steps that are both logically valid and relevant to the hypothesis. Instead, they tend to hallucinate invalid steps given the hypothesis. In this paper, we present a novel stepwise method, NLProofS (Natural Language Proof Search), which learns to generate relevant steps conditioning on the hypothesis. At the core of our approach, we train an independent verifier to check the validity of the proof steps to prevent hallucination. Instead of generating steps greedily, we search for proofs maximizing a global proof score judged by the verifier. NL-ProofS achieves state-of-the-art performance on EntailmentBank and RuleTaker. Specifically, it improves the correctness of predicted proofs from 27.7% to 33.3% in the distractor setting of EntailmentBank, demonstrating the effectiveness of NLProofS in generating challenging human-authored proofs. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Deductive Verification of Chain-of-Thought ReasoningZhan Ling, Yunhao Fang, Xuanlin Li, Zhiao Huang 等NeurIPS 2023 · 被引用 234 次
- Learning Deductive Reasoning from Synthetic Corpus based on Formal LogicTerufumi Morishita, Gaku Morio, Atsuki Yamaguchi, Yasuhiro SogawaICML 2023 · 被引用 45 次
- Entailer: Answering Questions with Faithful and Truthful Chains of ReasoningOyvind Tafjord, Bhavana Dalvi Mishra, Peter ClarkEMNLP 2022 · 被引用 28 次
- RECKONING: Reasoning through Dynamic Knowledge EncodingZeming Chen, Gail Weiss, Eric Mitchell, Asli Celikyilmaz 等NeurIPS 2023 · 被引用 21 次
- Taming Imperfect Process Verifiers: A Sampling Perspective on BacktrackingDhruv Rohatgi, Abhishek Shetty, Donya Saless, Yuchen Li 等ICLR 2026 · 被引用 15 次
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- A Benchmark for Systematic Generalization in Grounded Language UnderstandingLaura Ruis, Jacob Andreas, Marco Baroni, Diane Bouchacourt 等NeurIPS 2020 · 被引用 169 次
- Language Models of Code are Few-Shot Commonsense LearnersAman Madaan, Shuyan Zhou, Uri Alon, Yiming Yang 等EMNLP 2022 · 被引用 103 次
相关 Paper
- ProofInfer: Generating Proof via Iterative Hierarchical InferenceZichu Fei, Qi Zhang, Xin Zhou, Tao Gui 等EMNLP 2022
- NaturalProver: Grounded Mathematical Proof Generation with Language ModelsSean Welleck, Jiacheng Liu, Ximing Lu, Hannaneh Hajishirzi 等NeurIPS 2022 · 被引用 108 次
- RLET: A Reinforcement Learning Based Approach for Explainable QA with Entailment TreesTengxiao Liu, Qipeng Guo, Xiangkun Hu, Yue Zhang 等EMNLP 2022 · 被引用 8 次
- Natural Language Deduction with Incomplete InformationZayne Sprague, Kaj Bostrom, Swarat Chaudhuri, Greg DurrettEMNLP 2022 · 被引用 8 次
- Explaining Answers with Entailment TreesBhavana Dalvi, Peter Jansen, Oyvind Tafjord, Zhengnan Xie 等EMNLP 2021 · 被引用 6 次
