Learning to Explain: Datasets and Models for Identifying Valid Reasoning Chains in Multihop Question-Answering
Harsh Jhamtani, Peter Clark
Abstract
Despite the rapid progress in multihop question-answering (QA), models still have trouble explaining why an answer is correct, with limited explanation training data available to learn from. To address this, we introduce three explanation datasets in which explanations formed from corpus facts are annotated. Our first dataset, eQASC, contains over 98K explanation annotations for the multihop question answering dataset QASC, and is the first that annotates multiple candidate explanations for each answer. The second dataset eQASC-perturbed is constructed by crowd-sourcing perturbations (while preserving their validity) of a subset of explanations in QASC, to test consistency and generalization of explanation prediction models. The third dataset eOBQA is constructed by adding explanation annotations to the OBQA dataset to test generalization of models trained on eQASC. We show that this data can be used to significantly improve explanation quality (+14% absolute F1 over a strong retrieval baseline) using a BERT-based classifier, but still behind the upper bound, offering a new challenge for future research. We also explore a delexicalized chain representation in which repeated noun phrases are replaced by variables, thus turning them into generalized reasoning chains (for example: "X is a Y" AND "Y has Z" IMPLIES "X has Z"). We find that generalized chains maintain performance while also being more robust to certain perturbations. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d73fb1c7-2401-4174-8332-03bd63446bdcCited by top-tier papers20
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu et al.NeurIPS 2022 · 2,727 citations
- Selection-Inference: Exploiting Large Language Models for Interpretable Logical ReasoningAntonia Creswell, Murray Shanahan, Irina HigginsICLR 2023 · 110 citations
- GNN is a Counter? Revisiting GNN for Question AnsweringKuan Wang, Yuyu Zhang, Diyi Yang, Le Song et al.ICLR 2022 · 37 citations
- ExplaGraphs: An Explanation Graph Generation Task for Structured Commonsense ReasoningSwarnadeep Saha, Prateek Yadav, Lisa Bauer, Mohit BansalEMNLP 2021 · 27 citations
- LAMBADA: Backward Chaining for Automated Reasoning in Natural LanguageMehran Kazemi, Najoung Kim, Deepti Bhatia, Xin Xu et al.ACL 2023 · 27 citations
Builds on2
Related papers
- Triple-Fact Retriever: An explainable reasoning retrieval model for multi-hop QA problemChengmin Wu, Enrui Hu, Ke Zhan, Lan Luo et al.ICDE 2022 · 5 citations
- On the Challenges of Evaluating Compositional Explanations in Multi-Hop Inference: Relevance, Completeness, and Expert RatingsPeter Jansen, Kelly J. Smith, Dan Moreno, Huitzilin OrtizEMNLP 2021
- Explaining Answers with Entailment TreesBhavana Dalvi, Peter Jansen, Oyvind Tafjord, Zhengnan Xie et al.EMNLP 2021 · 6 citations
- Answering Questions by Meta-Reasoning over Multiple Chains of ThoughtOri Yoran, Tomer Wolfson, Ben Bogin, Uri Katz et al.EMNLP 2023 · 30 citations
- Explanations for CommonsenseQA: New Dataset and ModelsShourya Aggarwal, Divyanshu Mandowara, Vishwajeet Agrawal, Dinesh Khandelwal et al.ACL 2021
