The Sensitivity of Language Models and Humans to Winograd Schema Perturbations
Mostafa Abdou, Vinit Ravishankar, Maria Barrett, Yonatan Belinkov, Desmond Elliott, Anders Søgaard
Abstract
Large-scale pretrained language models are the major driving force behind recent improvements in performance on the Winograd Schema Challenge, a widely employed test of commonsense reasoning ability. We show, however, with a new diagnostic dataset, that these models are sensitive to linguistic perturbations of the Winograd examples that minimally affect human understanding. Our results highlight interesting differences between humans and language models: language models are more sensitive to number or gender alternations and synonym replacements than humans, and humans are more stable and consistent in their predictions, maintain a much higher absolute performance, and perform better on non-associative instances than associative ones. Overall, humans are correct more often than out-of-the-box models, and the models are sometimes right for the wrong reasons. Finally, we show that fine-tuning on a large, task-specific dataset can offer a solution to these issues.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Back to Square One: Artifact Detection, Training and Commonsense Disentanglement in the Winograd SchemaYanai Elazar, Hongming Zhang, Yoav Goldberg, Dan RothEMNLP 2021 · 25 citations
- Adversarial Attack against Cross-lingual Knowledge Graph AlignmentZeru Zhang, Zijie Zhang, Yang Zhou, Lingfei Wu et al.EMNLP 2021 · 10 citations
- BRAINTEASER: Lateral Thinking Puzzles for Large Language ModelsYifan Jiang, Filip Ilievski, Kaixin Ma, Zhivar SouratiEMNLP 2023 · 6 citations
- Transformers in the loop: Polarity in neural models of languageLisa Bylinina, Alexey TikhonovACL 2022
- A Semantic-based Method for Unsupervised Commonsense Question AnsweringYilin Niu, Fei Huang, Jiaming Liang, Wenkai Chen et al.ACL 2021
Builds on2
Related papers
- WinoLogic: A Zero-Shot Logic-based Diagnostic Dataset for Winograd Schema ChallengeWeinan He, Canming Huang, Yongmei Liu, Xiaodan ZhuEMNLP 2021 · 7 citations
- WinoWhy: A Deep Diagnosis of Essential Commonsense Knowledge for Answering Winograd Schema ChallengeHongming Zhang, Xinran Zhao, Yangqiu SongACL 2020 · 33 citations
- Winoground: Probing Vision and Language Models for Visio-Linguistic CompositionalityTristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh et al.CVPR 2022 · 179 citations
- Picturing Ambiguity: A Visual Twist on the Winograd Schema ChallengeBrendan Park, Madeline Janecek, Naser Ezzati-Jivan, Yifeng Li et al.ACL 2024
- Rethinking Pragmatics in Large Language Models: Towards Open-Ended Evaluation and Preference TuningShengguang Wu, Shusheng Yang, Zhenglun Chen, Qi SuEMNLP 2024 · 2 citations
