Shortcutted Commonsense: Data Spuriousness in Deep Learning of Commonsense Reasoning
Ruben Branco, António Branco, João António Rodrigues, João Ricardo Silva
摘要
Commonsense is a quintessential human capacity that has been a core challenge to Artificial Intelligence since its inception. Impressive results in Natural Language Processing tasks, including in commonsense reasoning, have consistently been achieved with Transformer neural language models, even matching or surpassing human performance in some benchmarks. Recently, some of these advances have been called into question: so called data artifacts in the training data have been made evident as spurious correlations and shallow shortcuts that in some cases are leveraging these outstanding results. In this paper we seek to further pursue this analysis into the realm of commonsense related language processing tasks. We undertake a study on different prominent benchmarks that involve commonsense reasoning, along a number of key stress experiments, thus seeking to gain insight on whether the models are learning transferable generalizations intrinsic to the problem at stake or just taking advantage of incidental shortcuts in the data items. The results obtained indicate that most datasets experimented with are problematic, with models resorting to non-robust features and appearing not to be learning and generalizing towards the overall tasks intended to be conveyed or exemplified by the datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- SCOTT: Self-Consistent Chain-of-Thought DistillationPeifeng Wang, Zhengyang Wang, Zheng Li, Yifan Gao 等ACL 2023 · 被引用 39 次
- Say What You Mean! Large Language Models Speak Too Positively about Negative Commonsense KnowledgeJiangjie Chen, Wei Shi, Ziquan Fu, Sijie Cheng 等ACL 2023 · 被引用 23 次
- PINTO: Faithful Language Reasoning Using Prompt-Generated RationalesPeifeng Wang, Aaron Chan, Filip Ilievski, Muhao Chen 等ICLR 2023 · 被引用 21 次
- BRAINTEASER: Lateral Thinking Puzzles for Large Language ModelsYifan Jiang, Filip Ilievski, Kaixin Ma, Zhivar SouratiEMNLP 2023 · 被引用 6 次
- Does Your Model Classify Entities Reasonably? Diagnosing and Mitigating Spurious Correlations in Entity TypingNan Xu, Fei Wang, Bangzheng Li, Mingtao Dong 等EMNLP 2022 · 被引用 6 次
它引用的顶会 Paper7
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 被引用 1,333 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- (Comet-) Atomic 2020: On Symbolic and Neural Commonsense Knowledge GraphsJena D. Hwang, Chandra Bhagavatula, Ronan Le Bras, Jeff Da 等AAAI 2021 · 被引用 458 次
- Evaluating Commonsense in Pre-Trained Language ModelsXuhui Zhou, Yue Zhang, Leyang Cui, Dandan HuangAAAI 2020 · 被引用 198 次
相关 Paper
- Do Vision-Language Transformers Exhibit Visual Commonsense? An Empirical Study of VCRZhenyang Li, Yangyang Guo, Kejie Wang, Xiaolin Chen 等ACM MM 2023 · 被引用 11 次
- UNICORN on RAINBOW: A Universal Commonsense Reasoning Model on a New Multitask BenchmarkNicholas Lourie, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2021 · 被引用 149 次
- A Systematic Investigation of Commonsense Knowledge in Large Language ModelsXiang Lorraine Li, Adhiguna Kuncoro, Jordan Hoffmann, Cyprien de Masson d'Autume 等EMNLP 2022 · 被引用 34 次
- RICA: Evaluating Robust Inference Capabilities Based on Commonsense AxiomsPei Zhou, Rahul Khanna, Seyeon Lee, Bill Yuchen Lin 等EMNLP 2021 · 被引用 28 次
- RussianSuperGLUE: A Russian Language Understanding Evaluation BenchmarkTatiana Shavrina, Alena Fenogenova, Anton A. Emelyanov, Denis Shevelev 等EMNLP 2020 · 被引用 11 次
