A Systematic Investigation of Commonsense Knowledge in Large Language Models
Xiang Lorraine Li, Adhiguna Kuncoro, Jordan Hoffmann, Cyprien de Masson d'Autume, Phil Blunsom, Aida Nematzadeh
摘要
Language models (LMs) trained on large amounts of data (e.g., Brown et al., 2020; Patwary et al., 2021) have shown impressive performance on many NLP tasks under the zeroshot and few-shot setup. Here we aim to better understand the extent to which such models learn commonsense knowledge -a critical component of many NLP applications. We conduct a systematic and rigorous zero-shot and few-shot commonsense evaluation of large pretrained LMs, where we: (i) carefully control for the LMs' ability to exploit potential surface cues and annotation artefacts, and (ii) account for variations in performance that arise from factors that are not related to commonsense knowledge. Our findings highlight the limitations of pre-trained LMs in acquiring commonsense knowledge without task-specific supervision; furthermore, using larger models or few-shot evaluation are insufficient to achieve human-level commonsense performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Adaptive Chameleon or Stubborn Sloth: Revealing the Behavior of Large Language Models in Knowledge ConflictsJian Xie, Kai Zhang, Jiangjie Chen, Renze Lou 等ICLR 2024 · 被引用 294 次
- OMNI: Open-endedness via Models of human Notions of InterestingnessJenny Zhang, Joel Lehman, Kenneth O. Stanley, Jeff CluneICLR 2024 · 被引用 59 次
- Tree-Planner: Efficient Close-loop Task Planning with Large Language ModelsMengkang Hu, Yao Mu, Xinmiao Yu, Mingyu Ding 等ICLR 2024 · 被引用 57 次
- Thought Cloning: Learning to Think while Acting by Imitating Human ThinkingShengran Hu, Jeff CluneNeurIPS 2023 · 被引用 49 次
- Large Language Models are Temporal and Causal Reasoners for Video Question AnsweringDohwan Ko, Ji Soo Lee, Woo-Young Kang, Byungseok Roh 等EMNLP 2023 · 被引用 30 次
它引用的顶会 Paper14
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe 等EMNLP 2022 · 被引用 634 次
- Abductive Commonsense ReasoningChandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi 等ICLR 2020 · 被引用 521 次
相关 Paper
- Back to Square One: Artifact Detection, Training and Commonsense Disentanglement in the Winograd SchemaYanai Elazar, Hongming Zhang, Yoav Goldberg, Dan RothEMNLP 2021 · 被引用 25 次
- Zero-Shot Commonsense Question Answering with Cloze Translation and Consistency OptimizationZi-Yi Dou, Nanyun PengAAAI 2022 · 被引用 29 次
- Prompting Language Models for Linguistic StructureTerra Blevins, Hila Gonen, Luke ZettlemoyerACL 2023 · 被引用 15 次
- RICA: Evaluating Robust Inference Capabilities Based on Commonsense AxiomsPei Zhou, Rahul Khanna, Seyeon Lee, Bill Yuchen Lin 等EMNLP 2021 · 被引用 28 次
- When Do You Need Billions of Words of Pretraining Data?Yian Zhang, Alex Warstadt, Xiaocheng Li, Samuel R. BowmanACL 2021
