Is Multihop QA in DiRe Condition? Measuring and Reducing Disconnected Reasoning
Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, Ashish Sabharwal
摘要
Has there been real progress in multi-hop question-answering? Models often exploit dataset artifacts to produce correct answers, without connecting information across multiple supporting facts. This limits our ability to measure true progress and defeats the purpose of building multi-hop QA datasets. We make three contributions towards addressing this. First, we formalize such undesirable behavior as disconnected reasoning across subsets of supporting facts. This allows developing a model-agnostic probe for measuring how much any model can cheat via disconnected reasoning. Second, using a notion of contrastive support sufficiency, we introduce an automatic transformation of existing datasets that reduces the amount of disconnected reasoning. Third, our experiments 1 suggest that there hasn't been much progress in multifact QA in the reading comprehension setting. For a recent large-scale model (XLNet), we show that only 18 points out of its answer F1 score of 72 on HotpotQA are obtained through multifact reasoning, roughly the same as that of a simpler RNN baseline. Our transformation substantially reduces disconnected reasoning (19 points in answer F1). It is complementary to adversarial approaches, yielding further reductions in conjunction.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Causal Walk: Debiasing Multi-Hop Fact Verification with Front-Door AdjustmentCongzhi Zhang, Linhai Zhang, Deyu ZhouAAAI 2024 · 被引用 32 次
- DISCO: Distilling Counterfactuals with Large Language ModelsZeming Chen, Qiyue Gao, Antoine Bosselut, Ashish Sabharwal 等ACL 2023 · 被引用 27 次
- Summarize-then-Answer: Generating Concise Explanations for Multi-hop Reading ComprehensionNaoya Inoue, Harsh Trivedi, Steven Sinha, Niranjan Balasubramanian 等EMNLP 2021 · 被引用 14 次
- Teaching Broad Reasoning Skills for Multi-Step QA by Generating Hard ContextsHarsh Trivedi, Niranjan Balasubramanian, Tushar Khot, Ashish SabharwalEMNLP 2022 · 被引用 6 次
- Counterfactual Multihop QA: A Cause-Effect Approach for Reducing Disconnected ReasoningWangzhen Guo, Qinkang Gong, Yanghui Rao, Hanjiang LaiACL 2023 · 被引用 4 次
它引用的顶会 Paper6
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 被引用 625 次
- QASC: A Dataset for Question Answering via Sentence CompositionTushar Khot, Peter Clark, Michal Guerquin, Peter Jansen 等AAAI 2020 · 被引用 387 次
- Learning to Retrieve Reasoning Paths over Wikipedia Graph for Question AnsweringAkari Asai, Kazuma Hashimoto, Hannaneh Hajishirzi, Richard Socher 等ICLR 2020 · 被引用 322 次
- Hierarchical Graph Network for Multi-hop Question AnsweringYuwei Fang, Siqi Sun, Zhe Gan, Rohit Pillai 等EMNLP 2020 · 被引用 157 次
- Select, Answer and Explain: Interpretable Multi-Hop Reading Comprehension over Multiple DocumentsMing Tu, Kevin Huang, Guangtao Wang, Jing Huang 等AAAI 2020 · 被引用 155 次
相关 Paper
- Triple-Fact Retriever: An explainable reasoning retrieval model for multi-hop QA problemChengmin Wu, Enrui Hu, Ke Zhan, Lan Luo 等ICDE 2022 · 被引用 5 次
- Transformer-XH: Multi-Evidence Reasoning with eXtra Hop AttentionChen Zhao, Chenyan Xiong, Corby Rosset, Xia Song 等ICLR 2020 · 被引用 120 次
- Modeling Multi-hop Question Answering as Single Sequence PredictionSemih Yavuz, Kazuma Hashimoto, Yingbo Zhou, Nitish Shirish Keskar 等ACL 2022 · 被引用 35 次
- Low-Resource Generation of Multi-hop Reasoning QuestionsJianxing Yu, Wei Liu, Shuang Qiu, Qinliang Su 等ACL 2020 · 被引用 11 次
- Robustifying Multi-hop QA through Pseudo-Evidentiality TrainingKyungjae Lee, Seung-won Hwang, Sang-eun Han, Dohyeon LeeACL 2021
