RLET: A Reinforcement Learning Based Approach for Explainable QA with Entailment Trees
Tengxiao Liu, Qipeng Guo, Xiangkun Hu, Yue Zhang, Xipeng Qiu, Zheng Zhang
摘要
Interpreting the reasoning process from questions to answers poses a challenge in approaching explainable QA. A recently proposed structured reasoning format, entailment tree, manages to offer explicit logical deductions with entailment steps in a tree structure. To generate entailment trees, prior single pass sequence-to-sequence models lack visible internal decision probability, while stepwise approaches are supervised with extracted single step data and cannot model the tree as a whole. In this work, we propose RLET, a Reinforcement Learning based Entailment Tree generation framework, which is trained utilising the cumulative signals across the whole tree. RLET iteratively performs single step reasoning with sentence selection and deduction generation modules, from which the training signal is accumulated across the tree with elaborately designed aligned reward function that is consistent with the evaluation. To the best of our knowledge, we are the first to introduce RL into the entailment tree generation task. Experiments on three settings of the En-tailmentBank dataset demonstrate the strength of using RL framework.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Explicit Planning Helps Language Models in Logical ReasoningHongyu Zhao, Kangrui Wang, Mo Yu, Hongyuan MeiEMNLP 2023 · 被引用 8 次
- Faithful Question Answering with Monte-Carlo PlanningRuixin Hong, Hongming Zhang, Hong Zhao, Dong Yu 等ACL 2023 · 被引用 7 次
- SEER: Facilitating Structured Reasoning and Explanation via Reinforcement LearningGuoxin Chen, Kexin Tang, Chao Yang, Fuying Ye 等ACL 2024 · 被引用 5 次
- An Entailment Tree Generation Approach for Multimodal Multi-Hop Question Answering with Mixture-of-Experts and Iterative Feedback MechanismQing Zhang, Haocheng Lv, Jie Liu, Zhiyun Chen 等ACM MM 2024 · 被引用 2 次
它引用的顶会 Paper10
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 被引用 394 次
- FaiRR: Faithful and Robust Deductive Reasoning over Natural LanguageSoumya Sanyal, Harman Singh, Xiang RenACL 2022 · 被引用 49 次
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 被引用 40 次
相关 Paper
- Explaining Answers with Entailment TreesBhavana Dalvi, Peter Jansen, Oyvind Tafjord, Zhengnan Xie 等EMNLP 2021 · 被引用 6 次
- TV-TREES: Multimodal Entailment Trees for Neuro-Symbolic Video ReasoningKate Sanders, Nathaniel Weir, Benjamin Van DurmeEMNLP 2024 · 被引用 3 次
- Neural Natural Logic Inference for Interpretable Question AnsweringJihao Shi, Xiao Ding, Li Du, Ting Liu 等EMNLP 2021 · 被引用 10 次
- Commonsense Video Question Answering through Video-Grounded Entailment Tree ReasoningHuabin Liu, Filip Ilievski, Cees G. M. SnoekCVPR 2025
- R5: Rule Discovery with Reinforced and Recurrent Relational ReasoningShengyao Lu, Bang Liu, Keith G. Mills, Shangling Jui 等ICLR 2022 · 被引用 10 次
