Is Multihop QA in DiRe Condition? Measuring and Reducing Disconnected Reasoning
Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, Ashish Sabharwal
Abstract
Has there been real progress in multi-hop question-answering? Models often exploit dataset artifacts to produce correct answers, without connecting information across multiple supporting facts. This limits our ability to measure true progress and defeats the purpose of building multi-hop QA datasets. We make three contributions towards addressing this. First, we formalize such undesirable behavior as disconnected reasoning across subsets of supporting facts. This allows developing a model-agnostic probe for measuring how much any model can cheat via disconnected reasoning. Second, using a notion of contrastive support sufficiency, we introduce an automatic transformation of existing datasets that reduces the amount of disconnected reasoning. Third, our experiments 1 suggest that there hasn't been much progress in multifact QA in the reading comprehension setting. For a recent large-scale model (XLNet), we show that only 18 points out of its answer F1 score of 72 on HotpotQA are obtained through multifact reasoning, roughly the same as that of a simpler RNN baseline. Our transformation substantially reduces disconnected reasoning (19 points in answer F1). It is complementary to adversarial approaches, yielding further reductions in conjunction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- Causal Walk: Debiasing Multi-Hop Fact Verification with Front-Door AdjustmentCongzhi Zhang, Linhai Zhang, Deyu ZhouAAAI 2024 · 32 citations
- DISCO: Distilling Counterfactuals with Large Language ModelsZeming Chen, Qiyue Gao, Antoine Bosselut, Ashish Sabharwal et al.ACL 2023 · 27 citations
- Summarize-then-Answer: Generating Concise Explanations for Multi-hop Reading ComprehensionNaoya Inoue, Harsh Trivedi, Steven Sinha, Niranjan Balasubramanian et al.EMNLP 2021 · 14 citations
- Teaching Broad Reasoning Skills for Multi-Step QA by Generating Hard ContextsHarsh Trivedi, Niranjan Balasubramanian, Tushar Khot, Ashish SabharwalEMNLP 2022 · 6 citations
- Counterfactual Multihop QA: A Cause-Effect Approach for Reducing Disconnected ReasoningWangzhen Guo, Qinkang Gong, Yanghui Rao, Hanjiang LaiACL 2023 · 4 citations
Builds on6
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 625 citations
- QASC: A Dataset for Question Answering via Sentence CompositionTushar Khot, Peter Clark, Michal Guerquin, Peter Jansen et al.AAAI 2020 · 387 citations
- Learning to Retrieve Reasoning Paths over Wikipedia Graph for Question AnsweringAkari Asai, Kazuma Hashimoto, Hannaneh Hajishirzi, Richard Socher et al.ICLR 2020 · 322 citations
- Hierarchical Graph Network for Multi-hop Question AnsweringYuwei Fang, Siqi Sun, Zhe Gan, Rohit Pillai et al.EMNLP 2020 · 157 citations
- Select, Answer and Explain: Interpretable Multi-Hop Reading Comprehension over Multiple DocumentsMing Tu, Kevin Huang, Guangtao Wang, Jing Huang et al.AAAI 2020 · 155 citations
Related papers
- Triple-Fact Retriever: An explainable reasoning retrieval model for multi-hop QA problemChengmin Wu, Enrui Hu, Ke Zhan, Lan Luo et al.ICDE 2022 · 5 citations
- Transformer-XH: Multi-Evidence Reasoning with eXtra Hop AttentionChen Zhao, Chenyan Xiong, Corby Rosset, Xia Song et al.ICLR 2020 · 120 citations
- Modeling Multi-hop Question Answering as Single Sequence PredictionSemih Yavuz, Kazuma Hashimoto, Yingbo Zhou, Nitish Shirish Keskar et al.ACL 2022 · 35 citations
- Low-Resource Generation of Multi-hop Reasoning QuestionsJianxing Yu, Wei Liu, Shuang Qiu, Qinliang Su et al.ACL 2020 · 11 citations
- Robustifying Multi-hop QA through Pseudo-Evidentiality TrainingKyungjae Lee, Seung-won Hwang, Sang-eun Han, Dohyeon LeeACL 2021
