Back to Square One: Artifact Detection, Training and Commonsense Disentanglement in the Winograd Schema
Yanai Elazar, Hongming Zhang, Yoav Goldberg, Dan Roth
摘要
The Winograd Schema (WS) has been proposed as a test for measuring commonsense capabilities of models. Recently, pre-trained language model-based approaches have boosted performance on some WS benchmarks but the source of improvement is still not clear. This paper suggests that the apparent progress on WS may not necessarily reflect progress in commonsense reasoning. To support this claim, we first show that the current evaluation method of WS is sub-optimal and propose a modification that uses twin sentences for evaluation. We also propose two new baselines that indicate the existence of artifacts in WS benchmarks. We then develop a method for evaluating WS-like sentences in a zero-shot setting to account for the commonsense reasoning abilities acquired during the pretraining and observe that popular language models perform randomly in this setting when using our more strict evaluation. We conclude that the observed progress is mostly due to the use of supervision in training WS models, which is not likely to successfully support all the required commonsense reasoning skills and knowledge. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Winoground: Probing Vision and Language Models for Visio-Linguistic CompositionalityTristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh 等CVPR 2022 · 被引用 179 次
- What's In My Big Data?Yanai Elazar, Akshita Bhagia, Ian Magnusson, Abhilasha Ravichander 等ICLR 2024 · 被引用 135 次
- CompA: Addressing the Gap in Compositional Reasoning in Audio-Language ModelsSreyan Ghosh, Ashish Seth, Sonal Kumar, Utkarsh Tyagi 等ICLR 2024 · 被引用 53 次
- Explanation Graph Generation via Pre-trained Language Models: An Empirical Study with Contrastive LearningSwarnadeep Saha, Prateek Yadav, Mohit BansalACL 2022 · 被引用 10 次
- Verification and Refinement of Natural Language Explanations through LLM-Symbolic Theorem ProvingXin Quan, Marco Valentino, Louise A. Dennis, André FreitasEMNLP 2024 · 被引用 9 次
它引用的顶会 Paper9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 被引用 625 次
相关 Paper
- The Sensitivity of Language Models and Humans to Winograd Schema PerturbationsMostafa Abdou, Vinit Ravishankar, Maria Barrett, Yonatan Belinkov 等ACL 2020 · 被引用 1 次
- WinoLogic: A Zero-Shot Logic-based Diagnostic Dataset for Winograd Schema ChallengeWeinan He, Canming Huang, Yongmei Liu, Xiaodan ZhuEMNLP 2021 · 被引用 7 次
- WinoWhy: A Deep Diagnosis of Essential Commonsense Knowledge for Answering Winograd Schema ChallengeHongming Zhang, Xinran Zhao, Yangqiu SongACL 2020 · 被引用 33 次
- A Systematic Investigation of Commonsense Knowledge in Large Language ModelsXiang Lorraine Li, Adhiguna Kuncoro, Jordan Hoffmann, Cyprien de Masson d'Autume 等EMNLP 2022 · 被引用 34 次
- Zero-Shot Commonsense Question Answering with Cloze Translation and Consistency OptimizationZi-Yi Dou, Nanyun PengAAAI 2022 · 被引用 29 次
