WinoWhy: A Deep Diagnosis of Essential Commonsense Knowledge for Answering Winograd Schema Challenge
Hongming Zhang, Xinran Zhao, Yangqiu Song
Abstract
In this paper, we present the first comprehensive categorization of essential commonsense knowledge for answering the Winograd Schema Challenge (WSC). For each of the questions, we invite annotators to first provide reasons for making correct decisions and then categorize them into six major knowledge categories. By doing so, we better understand the limitation of existing methods (i.e., what kind of knowledge cannot be effectively represented or inferred with existing methods) and shed some light on the commonsense knowledge that we need to acquire in the future for better commonsense reasoning. Moreover, to investigate whether current WSC models can understand the commonsense or they simply solve the WSC questions based on the statistical bias of the dataset, we leverage the collected reasons to develop a new task called WinoWhy, which requires models to distinguish plausible reasons from very similar but wrong reasons for all WSC questions. Experimental results prove that even though pre-trained language representation models have achieved promising progress on the original WSC dataset, they are still struggling at WinoWhy. Further experiments show that even though supervised models can achieve better performance, the performance of these models can be sensitive to the dataset distribution. WinoWhy and all codes are available at: https://github.com/ HKUST-KnowComp/WinoWhy .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0708b7b9-2ff5-4dbc-93f8-74ec08e65302Cited by top-tier papers10
- Back to Square One: Artifact Detection, Training and Commonsense Disentanglement in the Winograd SchemaYanai Elazar, Hongming Zhang, Yoav Goldberg, Dan RothEMNLP 2021 · 25 citations
- More is not always better? Enhancing Many-Shot In-Context Learning with Differentiated and Reweighting ObjectivesXiaoqing Zhang, Ang Lv, Yuhan Liu, Flood Sung et al.ACL 2025 · 9 citations
- WinoLogic: A Zero-Shot Logic-based Diagnostic Dataset for Winograd Schema ChallengeWeinan He, Canming Huang, Yongmei Liu, Xiaodan ZhuEMNLP 2021 · 7 citations
- Hop, Union, Generate: Explainable Multi-hop Reasoning without Rationale SupervisionWenting Zhao, Justin T. Chiu, Claire Cardie, Alexander M. RushEMNLP 2023 · 4 citations
- ALERT: Adapt Language Models to Reasoning TasksPing Yu, Tianlu Wang, Olga Golovneva, Badr AlKhamissi et al.ACL 2023 · 4 citations
Builds on2
Related papers
- The Sensitivity of Language Models and Humans to Winograd Schema PerturbationsMostafa Abdou, Vinit Ravishankar, Maria Barrett, Yonatan Belinkov et al.ACL 2020 · 1 citation
- Why is Winoground Hard? Investigating Failures in Visuolinguistic CompositionalityAnuj Diwan, Layne Berry, Eunsol Choi, David Harwath et al.EMNLP 2022 · 15 citations
- Winoground: Probing Vision and Language Models for Visio-Linguistic CompositionalityTristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh et al.CVPR 2022 · 179 citations
- WikiWhy: Answering and Explaining Cause-and-Effect QuestionsMatthew Ho, Aditya Sharma, Justin Chang, Michael Saxon et al.ICLR 2023 · 8 citations
- FOCUS: Evaluating Pre-trained Vision-Language Models on Underspecification ReasoningKankan Zhou, Eason Lai, Kyriakos Mouratidis, Jing JiangACL 2025
