CREPE: Open-Domain Question Answering with False Presuppositions
Xinyan Yu, Sewon Min, Luke Zettlemoyer, Hannaneh Hajishirzi
摘要
When asking about unfamiliar topics, information seeking users often pose questions with false presuppositions. Most existing question answering (QA) datasets, in contrast, assume all questions have well defined answers. We introduce CREPE, a QA dataset containing a natural distribution of presupposition failures from online information-seeking forums. We find that 25% of questions contain false presuppositions, and provide annotations for these presuppositions and their corrections. Through extensive baseline experiments, we show that adaptations of existing open-domain QA models can find presuppositions moderately well, but struggle when predicting whether a presupposition is factually correct. This is in large part due to difficulty in retrieving relevant evidence passages from a large text corpus. CREPE provides a benchmark to study question answering in the wild, and our analyses provide avenues for future work in better modeling and further studying the task. 1 Question: If there's an equal and opposite reaction for everything, how does any action happen? Isn't it balanced out by the opposite reaction? False presupposition: The equal and opposite reaction applies to the same object. Correction: Based on Newton's Law of Motion, the equal and opposite reaction applies to the other object. Only forces that are applied to the same object would be cancelled out. Newton's laws of motion From Wikipedia, the free encyclopedia Inputs given to the human raters Question: Why do prosecuters/courts seek/sentence prison time greater than the expected lifespan of the offender (i.e. 150 years in prison)? Why not simply sentence those criminals to 'life' in prison instead? Comment: Sentencing options are written into state laws. Life in prison is different in state laws than 150 years. Some of it comes into play with the "cruel and unusual punishment" clause in the Constitution too. Life in prison may not be "cruel and unusual" for a murder sentence, but it might be for, say, child sex trafficking. But if you trafficked 10 kids and the sentence is 15 years for each one, you get an effective life sentence that will also stand up, Constitutionally, against a "cruel and unusual punishment" defense. Outputs human raters rate Reference Presupposition: It does not make sense to sentence a person to 150 years in prison if they can't live that long anyways, prosecutors should use the life in prison sentence instead. Correction: The defendant can argue the life in prison sentence as cruel and unusual, so the actual year sentence is better to give than the alternative. GOLD-COMMENT track, Dedicated Presupposition: Penalties should be able to be sentenced to life in prison. Correction: Life in prison is different in state laws than 150 years in prison. GOLD-COMMENT track, Unified Presupposition: If a criminal is sentenced to life in prison, they should be sentenced to life in prison. Correction: It is not the case that if a criminal is sentenced to life in prison, they should be sentenced to life in prison. Main, Dedicated Presupposition: Penalties should be able to be imposed on criminals for life. Correction: The longer the sentence, the more likely the prosecution will seek to sentence the offender to life in prison. Main, Unified Presupposition: Prosecutor's should seek prison time greater than the expected lifespan of the offender. Correction: It is not the case that prosecutor's should seek prison time greater than the expected lifespan of the offender. Table 13 : An example of the input and the output human raters are given for the human evaluation of the writing subtask. Note that human raters are not given which output is a reference or from which system. Inputs given to the human raters Question: Why did scientists in the 1970s think that there was going to be a new ice age soon? Comment: They didn't. Between 1965 and 1979, there was 7 papers talking about global cooling (not ice age and not necessarily soon). During the same period there was 44 papers about global warming. The media just liked the sensationalism, so there was some news article and a front page on the Times Magazine. They started with a minority of scientist talking about global cooling in a time period when there was still a lot of unknown in climate science and changed that to Scientific consensus that an Ice Age is coming soon.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- Can Language Models Solve Graph Problems in Natural Language?Heng Wang, Shangbin Feng, Tianxing He, Zhaoxuan Tan 等NeurIPS 2023 · 被引用 420 次
- FLASK: Fine-grained Language Model Evaluation based on Alignment Skill SetsSeonghyeon Ye, Doyoung Kim, Sungdong Kim, Hyeonbin Hwang 等ICLR 2024 · 被引用 176 次
- Task Contamination: Language Models May Not Be Few-Shot AnymoreChangmao Li, Jeffrey FlaniganAAAI 2024 · 被引用 138 次
- Towards Understanding Factual Knowledge of Large Language ModelsXuming Hu, Junzhe Chen, Xiaochuan Li, Yufei Guo 等ICLR 2024 · 被引用 21 次
- Cancer-Myth: Evaluating Large Language Models on Patient Questions with False PresuppositionsWang Zhu, Tianqi Chen, Xinyan Yu, Ching Ying Lin 等ICLR 2026 · 被引用 15 次
它引用的顶会 Paper9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Autoregressive Entity RetrievalNicola De Cao, Gautier Izacard, Sebastian Riedel, Fabio PetroniICLR 2021 · 被引用 200 次
- AmbigQA: Answering Ambiguous Open-domain QuestionsSewon Min, Julian Michael, Hannaneh Hajishirzi, Luke ZettlemoyerEMNLP 2020 · 被引用 162 次
- How Do We Answer Complex Questions: Discourse Structure of Long-form AnswersFangyuan Xu, Junyi Jessy Li, Eunsol ChoiACL 2022 · 被引用 25 次
- Social Bias Frames: Reasoning about Social and Power Implications of LanguageMaarten Sap, Saadia Gabriel, Lianhui Qin, Dan Jurafsky 等ACL 2020 · 被引用 16 次
相关 Paper
- (QA)²: Question Answering with Questionable AssumptionsNajoung Kim, Phu Mon Htut, Samuel R. Bowman, Jackson PettyACL 2023 · 被引用 2 次
- Which Linguist Invented the Lightbulb? Presupposition Verification for Question-AnsweringNajoung Kim, Ellie Pavlick, Burcu Karagol Ayan, Deepak RamachandranACL 2021
- IfQA: A Dataset for Open-domain Question Answering under Counterfactual PresuppositionsWenhao Yu, Meng Jiang, Peter Clark, Ashish SabharwalEMNLP 2023 · 被引用 6 次
- Pragmatic Reasoning Unlocks Quantifier Semantics for Foundation ModelsYiyuan Li, Rakesh R. Menon, Sayan Ghosh, Shashank SrivastavaEMNLP 2023
- Won't Get Fooled Again: Answering Questions with False PremisesShengding Hu, Yifan Luo, Huadong Wang, Xingyi Cheng 等ACL 2023 · 被引用 5 次
