ProtoQA: A Question Answering Dataset for Prototypical Common-Sense Reasoning
Michael Boratko, Xiang Li, Tim O'Gorman, Rajarshi Das, Dan Le, Andrew McCallum
Abstract
Given questions regarding some prototypical situation -such as Name something that people usually do before they leave the house for work? -a human can easily answer them via acquired experiences. There can be multiple right answers for such questions, with some more common for a situation than others. This paper introduces a new question answering dataset for training and evaluating common sense reasoning capabilities of artificial intelligence systems in such prototypical situations. The training set is gathered from an existing set of questions played in a longrunning international game show -FAMILY-FEUD. The hidden evaluation set is created by gathering answers for each question from 100 crowd-workers. We also propose a generative evaluation task where a model has to output a ranked list of answers, ideally covering all prototypical answers for a question. After presenting multiple competitive baseline models, we find that human performance still exceeds model scores on all evaluation metrics with a meaningful gap, supporting the challenging nature of the task. * Equal contribution. (i) Name something that people usually do before they leave for work? Ask 100 crowd-workers + manual clustering
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 25f8b94a-ebe2-4fa2-a2a1-f56fa39b7500Cited by top-tier papers16
- A Systematic Investigation of Commonsense Knowledge in Large Language ModelsXiang Lorraine Li, Adhiguna Kuncoro, Jordan Hoffmann, Cyprien de Masson d'Autume et al.EMNLP 2022 · 34 citations
- ExplaGraphs: An Explanation Graph Generation Task for Structured Commonsense ReasoningSwarnadeep Saha, Prateek Yadav, Lisa Bauer, Mohit BansalEMNLP 2021 · 27 citations
- : Visualizing and Understanding Commonsense Reasoning Capabilities of Natural Language ModelsXingbo Wang, Renfei Huang, Zhihua Jin, Tianqing Fang et al.IEEE VIS 2023 · 18 citations
- NormBank: A Knowledge Bank of Situational Social NormsCaleb Ziems, Jane Dwivedi-Yu, Yi-Chia Wang, Alon Y. Halevy et al.ACL 2023 · 15 citations
- Learning to Initialize: Can Meta Learning Improve Cross-task Generalization in Prompt Tuning?Chengwei Qin, Shafiq R. Joty, Qian Li, Ruochen ZhaoACL 2023 · 8 citations
Builds on10
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 625 citations
- Abductive Commonsense ReasoningChandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi et al.ICLR 2020 · 521 citations
Related papers
- CRoW: Benchmarking Commonsense Reasoning in Real-World TasksMete Ismayilzada, Debjit Paul, Syrielle Montariol, Mor Geva et al.EMNLP 2023 · 3 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- WikiWhy: Answering and Explaining Cause-and-Effect QuestionsMatthew Ho, Aditya Sharma, Justin Chang, Michael Saxon et al.ICLR 2023 · 8 citations
- Getting Closer to AI Complete Question Answering: A Set of Prerequisite Real TasksAnna Rogers, Olga Kovaleva, Matthew Downey, Anna RumshiskyAAAI 2020 · 141 citations
- How much coffee was consumed during EMNLP 2019? Fermi Problems: A New Reasoning Challenge for AIAshwin Kalyan, Abhinav Kumar, Arjun Chandrasekaran, Ashish Sabharwal et al.EMNLP 2021 · 1 citation
