Evaluating Commonsense in Pre-Trained Language Models
Xuhui Zhou, Yue Zhang, Leyang Cui, Dandan Huang
Abstract
Contextualized representations trained over large raw text data have given remarkable improvements for NLP tasks including question answering and reading comprehension. There have been works showing that syntactic, semantic and word sense knowledge are contained in such representations, which explains why they benefit such tasks. However, relatively little work has been done investigating commonsense knowledge contained in contextualized representations, which is crucial for human question answering and reading comprehension. We study the commonsense ability of GPT, BERT, XLNet, and RoBERTa by testing them on seven challenging benchmarks, finding that language modeling and its variants are effective objectives for promoting models' commonsense ability while bi-directional context and larger training set are bonuses. We additionally find that current models do poorly on tasks require more necessary inference steps. Finally, we test the robustness of models by making dual test cases, which are correlated so that the correct prediction of one sample should lead to correct prediction of the other. Interestingly, the models show confusion on these test cases, which suggests that they learn commonsense at the surface rather than the deep level. We release a test set, named CATs publicly, for future research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e1ef0c9b-8701-41e7-bf1c-396121809433Cited by top-tier papers30
- Factuality Enhanced Language Models for Open-Ended Text GenerationNayeon Lee, Wei Ping, Peng Xu, Mostofa Patwary et al.NeurIPS 2022 · 318 citations
- LIFT: Language-Interfaced Fine-Tuning for Non-language Machine Learning TasksTuan Dinh, Yuchen Zeng, Ruisu Zhang, Ziqian Lin et al.NeurIPS 2022 · 222 citations
- Knowledge-driven Data Construction for Zero-shot Evaluation in Commonsense Question AnsweringKaixin Ma, Filip Ilievski, Jonathan Francis, Yonatan Bisk et al.AAAI 2021 · 100 citations
- Benchmarking Knowledge-Enhanced Commonsense Question Answering via Knowledge-to-Text TransformationNing Bian, Xianpei Han, Bo Chen, Le SunAAAI 2021 · 49 citations
- Do LLMs Understand Social Knowledge? Evaluating the Sociability of Large Language Models with SocKET BenchmarkMinje Choi, Jiaxin Pei, Sagar Kumar, Chang Shu et al.EMNLP 2023 · 36 citations
Related papers
- When Do You Need Billions of Words of Pretraining Data?Yian Zhang, Alex Warstadt, Xiaocheng Li, Samuel R. BowmanACL 2021
- Intermediate-Task Transfer Learning with Pretrained Language Models: When and Why Does It Work?Yada Pruksachatkun, Jason Phang, Haokun Liu, Phu Mon Htut et al.ACL 2020 · 168 citations
- A Systematic Investigation of Commonsense Knowledge in Large Language ModelsXiang Lorraine Li, Adhiguna Kuncoro, Jordan Hoffmann, Cyprien de Masson d'Autume et al.EMNLP 2022 · 34 citations
- Compositional and Lexical Semantics in RoBERTa, BERT and DistilBERT: A Case Study on CoQAIeva Staliunaite, Ignacio IacobacciEMNLP 2020 · 2 citations
- Semantics-Aware BERT for Language UnderstandingZhuosheng Zhang, Yuwei Wu, Hai Zhao, Zuchao Li et al.AAAI 2020 · 396 citations
