WinoLogic: A Zero-Shot Logic-based Diagnostic Dataset for Winograd Schema Challenge
Weinan He, Canming Huang, Yongmei Liu, Xiaodan Zhu
摘要
The recent success of neural language models (NLMs) on the Winograd Schema Challenge has called for further investigation of the commonsense reasoning ability of these models. Previous diagnostic datasets rely on crowd-sourcing which fails to provide coherent commonsense crucial for solving WSC problems. To better evaluate NLMs, we propose a logic-based framework that focuses on highquality commonsense knowledge. Specifically, we identify and collect formal knowledge formulas verified by theorem provers and translate such formulas into natural language sentences. Based on these true knowledge sentences, adversarial false ones are generated. We propose a new dataset named WINOLOGIC with these sentences. Given a problem in WINOLOGIC, NLMs need to decide whether the plausible knowledge sentences could correctly solve the corresponding WSC problems in a zero-shot setting. We also ask human annotators to validate WINOLOGIC to ensure it is humanagreeable. Experiments show that NLMs still struggle to comprehend commonsense knowledge as humans do, indicating that their reasoning ability could have been overestimated.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- FOLIO: Natural Language Reasoning with First-Order LogicSimeng Han, Hailey Schoelkopf, Yilun Zhao, Zhenting Qi 等EMNLP 2024 · 被引用 18 次
- LLM-based Typed Hyperresolution for Commonsense Reasoning with Knowledge BasesArmin Toroghi, Ali Pesaranghader, Tanmana Sadhu, Scott SannerICLR 2025
它引用的顶会 Paper2
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- WinoWhy: A Deep Diagnosis of Essential Commonsense Knowledge for Answering Winograd Schema ChallengeHongming Zhang, Xinran Zhao, Yangqiu SongACL 2020 · 被引用 33 次
相关 Paper
- The Sensitivity of Language Models and Humans to Winograd Schema PerturbationsMostafa Abdou, Vinit Ravishankar, Maria Barrett, Yonatan Belinkov 等ACL 2020 · 被引用 1 次
- Back to Square One: Artifact Detection, Training and Commonsense Disentanglement in the Winograd SchemaYanai Elazar, Hongming Zhang, Yoav Goldberg, Dan RothEMNLP 2021 · 被引用 25 次
- RULEBREAKERS: Challenging LLMs at the Crossroads between Formal Logic and Human-like ReasoningJason Chan, Robert J. Gaizauskas, Zhixue ZhaoICML 2025
- Wino-X: Multilingual Winograd Schemas for Commonsense Reasoning and Coreference ResolutionDenis Emelin, Rico SennrichEMNLP 2021 · 被引用 12 次
- Winoground: Probing Vision and Language Models for Visio-Linguistic CompositionalityTristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh 等CVPR 2022 · 被引用 179 次
