(QA)²: Question Answering with Questionable Assumptions
Najoung Kim, Phu Mon Htut, Samuel R. Bowman, Jackson Petty
摘要
Naturally occurring information-seeking questions often contain questionable assumptions-assumptions that are false or unverifiable. Questions containing questionable assumptions are challenging because they require a distinct answer strategy that deviates from typical answers for information-seeking questions. For instance, the question When did Marie Curie discover Uranium? cannot be answered as a typical when question without addressing the false assumption Marie Curie discovered Uranium. In this work, we propose (QA) 2 (Question Answering with Questionable Assumptions), an open-domain evaluation dataset consisting of naturally occurring search engine queries that may or may not contain questionable assumptions. To be successful on (QA) 2 , systems must be able to detect questionable assumptions and also be able to produce adequate responses for both typical information-seeking questions and ones with questionable assumptions. Through human rater acceptability on end-to-end QA with (QA) 2 , we find that current models do struggle with handling questionable assumptions, leaving substantial headroom for progress. * Equal contribution, corresponding authors ∆ Work partly done at NYU before joining BU. δ Work done at NYU before joining Amazon. 1 We use the term questionable assumptions instead of presupposition failure to capture failures of both true presup-
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Measuring Chain of Thought Faithfulness by Unlearning Reasoning StepsMartin Tutek, Fateme Hashemi Chaleshtori, Ana Marasovic, Yonatan BelinkovEMNLP 2025 · 被引用 37 次
- Towards Understanding Factual Knowledge of Large Language ModelsXuming Hu, Junzhe Chen, Xiaochuan Li, Yufei Guo 等ICLR 2024 · 被引用 21 次
- EvolveBench: A Comprehensive Benchmark for Assessing Temporal Awareness in LLMs on Evolving KnowledgeZhiyuan Zhu, Yusheng Liao, Zhe Chen, Yuhao Wang 等ACL 2025 · 被引用 10 次
- No Questions are Stupid, but some are Poorly Posed: Understanding Poorly-Posed Information-Seeking QuestionsNeha Srikanth, Rachel Rudinger, Jordan Lee Boyd-GraberACL 2025 · 被引用 6 次
- Whispers that Shake Foundations: Analyzing and Mitigating False Premise Hallucinations in Large Language ModelsHongbang Yuan, Pengfei Cao, Zhuoran Jin, Yubo Chen 等EMNLP 2024 · 被引用 5 次
它引用的顶会 Paper11
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat 等ICML 2020 · 被引用 2,937 次
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach 等ICLR 2022 · 被引用 1,976 次
- Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context LearningHaokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta 等NeurIPS 2022 · 被引用 1,483 次
- Decomposed Prompting: A Modular Approach for Solving Complex TasksTushar Khot, Harsh Trivedi, Matthew Finlayson, Yao Fu 等ICLR 2023 · 被引用 94 次
相关 Paper
- Which Linguist Invented the Lightbulb? Presupposition Verification for Question-AnsweringNajoung Kim, Ellie Pavlick, Burcu Karagol Ayan, Deepak RamachandranACL 2021
- CREPE: Open-Domain Question Answering with False PresuppositionsXinyan Yu, Sewon Min, Luke Zettlemoyer, Hannaneh HajishirziACL 2023 · 被引用 13 次
- Identifying and Answering Questions with False Assumptions: An Interpretable ApproachZijie Wang, Eduardo BlancoEMNLP 2025
- Challenges in Information-Seeking QA: Unanswerable Questions and Paragraph RetrievalAkari Asai, Eunsol ChoiACL 2021
- IfQA: A Dataset for Open-domain Question Answering under Counterfactual PresuppositionsWenhao Yu, Meng Jiang, Peter Clark, Ashish SabharwalEMNLP 2023 · 被引用 6 次
