Evaluating Commonsense in Pre-Trained Language Models
Xuhui Zhou, Yue Zhang, Leyang Cui, Dandan Huang
摘要
Contextualized representations trained over large raw text data have given remarkable improvements for NLP tasks including question answering and reading comprehension. There have been works showing that syntactic, semantic and word sense knowledge are contained in such representations, which explains why they benefit such tasks. However, relatively little work has been done investigating commonsense knowledge contained in contextualized representations, which is crucial for human question answering and reading comprehension. We study the commonsense ability of GPT, BERT, XLNet, and RoBERTa by testing them on seven challenging benchmarks, finding that language modeling and its variants are effective objectives for promoting models' commonsense ability while bi-directional context and larger training set are bonuses. We additionally find that current models do poorly on tasks require more necessary inference steps. Finally, we test the robustness of models by making dual test cases, which are correlated so that the correct prediction of one sample should lead to correct prediction of the other. Interestingly, the models show confusion on these test cases, which suggests that they learn commonsense at the surface rather than the deep level. We release a test set, named CATs publicly, for future research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper30
- Factuality Enhanced Language Models for Open-Ended Text GenerationNayeon Lee, Wei Ping, Peng Xu, Mostofa Patwary 等NeurIPS 2022 · 被引用 318 次
- LIFT: Language-Interfaced Fine-Tuning for Non-language Machine Learning TasksTuan Dinh, Yuchen Zeng, Ruisu Zhang, Ziqian Lin 等NeurIPS 2022 · 被引用 222 次
- Knowledge-driven Data Construction for Zero-shot Evaluation in Commonsense Question AnsweringKaixin Ma, Filip Ilievski, Jonathan Francis, Yonatan Bisk 等AAAI 2021 · 被引用 100 次
- Benchmarking Knowledge-Enhanced Commonsense Question Answering via Knowledge-to-Text TransformationNing Bian, Xianpei Han, Bo Chen, Le SunAAAI 2021 · 被引用 49 次
- Do LLMs Understand Social Knowledge? Evaluating the Sociability of Large Language Models with SocKET BenchmarkMinje Choi, Jiaxin Pei, Sagar Kumar, Chang Shu 等EMNLP 2023 · 被引用 36 次
相关 Paper
- When Do You Need Billions of Words of Pretraining Data?Yian Zhang, Alex Warstadt, Xiaocheng Li, Samuel R. BowmanACL 2021
- Intermediate-Task Transfer Learning with Pretrained Language Models: When and Why Does It Work?Yada Pruksachatkun, Jason Phang, Haokun Liu, Phu Mon Htut 等ACL 2020 · 被引用 168 次
- A Systematic Investigation of Commonsense Knowledge in Large Language ModelsXiang Lorraine Li, Adhiguna Kuncoro, Jordan Hoffmann, Cyprien de Masson d'Autume 等EMNLP 2022 · 被引用 34 次
- Compositional and Lexical Semantics in RoBERTa, BERT and DistilBERT: A Case Study on CoQAIeva Staliunaite, Ignacio IacobacciEMNLP 2020 · 被引用 2 次
- Semantics-Aware BERT for Language UnderstandingZhuosheng Zhang, Yuwei Wu, Hai Zhao, Zuchao Li 等AAAI 2020 · 被引用 396 次
