Shortcutted Commonsense: Data Spuriousness in Deep Learning of Commonsense Reasoning
Ruben Branco, António Branco, João António Rodrigues, João Ricardo Silva
Abstract
Commonsense is a quintessential human capacity that has been a core challenge to Artificial Intelligence since its inception. Impressive results in Natural Language Processing tasks, including in commonsense reasoning, have consistently been achieved with Transformer neural language models, even matching or surpassing human performance in some benchmarks. Recently, some of these advances have been called into question: so called data artifacts in the training data have been made evident as spurious correlations and shallow shortcuts that in some cases are leveraging these outstanding results. In this paper we seek to further pursue this analysis into the realm of commonsense related language processing tasks. We undertake a study on different prominent benchmarks that involve commonsense reasoning, along a number of key stress experiments, thus seeking to gain insight on whether the models are learning transferable generalizations intrinsic to the problem at stake or just taking advantage of incidental shortcuts in the data items. The results obtained indicate that most datasets experimented with are problematic, with models resorting to non-robust features and appearing not to be learning and generalizing towards the overall tasks intended to be conveyed or exemplified by the datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c5dd2a89-7c0f-4fd6-abb8-d5d16446b10fCited by top-tier papers8
- SCOTT: Self-Consistent Chain-of-Thought DistillationPeifeng Wang, Zhengyang Wang, Zheng Li, Yifan Gao et al.ACL 2023 · 39 citations
- Say What You Mean! Large Language Models Speak Too Positively about Negative Commonsense KnowledgeJiangjie Chen, Wei Shi, Ziquan Fu, Sijie Cheng et al.ACL 2023 · 23 citations
- PINTO: Faithful Language Reasoning Using Prompt-Generated RationalesPeifeng Wang, Aaron Chan, Filip Ilievski, Muhao Chen et al.ICLR 2023 · 21 citations
- BRAINTEASER: Lateral Thinking Puzzles for Large Language ModelsYifan Jiang, Filip Ilievski, Kaixin Ma, Zhivar SouratiEMNLP 2023 · 6 citations
- Does Your Model Classify Entities Reasonably? Diagnosing and Mitigating Spurious Correlations in Entity TypingNan Xu, Fei Wang, Bangzheng Li, Mingtao Dong et al.EMNLP 2022 · 6 citations
Builds on7
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 1,333 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- (Comet-) Atomic 2020: On Symbolic and Neural Commonsense Knowledge GraphsJena D. Hwang, Chandra Bhagavatula, Ronan Le Bras, Jeff Da et al.AAAI 2021 · 458 citations
- Evaluating Commonsense in Pre-Trained Language ModelsXuhui Zhou, Yue Zhang, Leyang Cui, Dandan HuangAAAI 2020 · 198 citations
Related papers
- Do Vision-Language Transformers Exhibit Visual Commonsense? An Empirical Study of VCRZhenyang Li, Yangyang Guo, Kejie Wang, Xiaolin Chen et al.ACM MM 2023 · 11 citations
- UNICORN on RAINBOW: A Universal Commonsense Reasoning Model on a New Multitask BenchmarkNicholas Lourie, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2021 · 149 citations
- A Systematic Investigation of Commonsense Knowledge in Large Language ModelsXiang Lorraine Li, Adhiguna Kuncoro, Jordan Hoffmann, Cyprien de Masson d'Autume et al.EMNLP 2022 · 34 citations
- RICA: Evaluating Robust Inference Capabilities Based on Commonsense AxiomsPei Zhou, Rahul Khanna, Seyeon Lee, Bill Yuchen Lin et al.EMNLP 2021 · 28 citations
- RussianSuperGLUE: A Russian Language Understanding Evaluation BenchmarkTatiana Shavrina, Alena Fenogenova, Anton A. Emelyanov, Denis Shevelev et al.EMNLP 2020 · 11 citations
