Knowledge Rumination for Pre-trained Language Models
Yunzhi Yao, Peng Wang, Shengyu Mao, Chuanqi Tan, Fei Huang, Huajun Chen, Ningyu Zhang
摘要
Previous studies have revealed that vanilla pre-trained language models (PLMs) lack the capacity to handle knowledge-intensive NLP tasks alone; thus, several works have attempted to integrate external knowledge into PLMs. However, despite the promising outcome, we empirically observe that PLMs may have already encoded rich knowledge in their pre-trained parameters but fail to fully utilize them when applying them to knowledgeintensive tasks. In this paper, we propose a new paradigm dubbed Knowledge Rumination to help the pre-trained language model utilize that related latent knowledge without retrieving it from the external corpus. By simply adding a prompt like "As far as I know" to the PLMs, we try to review related latent knowledge and inject them back into the model for knowledge consolidation. We apply the proposed knowledge rumination to various language models, including RoBERTa, De-BERTa, and GPT-3. Experimental results on six commonsense reasoning tasks and GLUE benchmarks demonstrate the effectiveness of our proposed approach, which proves that the knowledge stored in PLMs can be better exploited to enhance performance 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Editing Large Language Models: Problems, Methods, and OpportunitiesYunzhi Yao, Peng Wang, Bozhong Tian, Siyuan Cheng 等EMNLP 2023 · 被引用 83 次
- Rule or Story, Which is a Better Commonsense Expression for Talking with Large Language Models?Ning Bian, Xianpei Han, Hongyu Lin, Yaojie Lu 等ACL 2024 · 被引用 1 次
- Why and How LLMs Benefit from Knowledge Introspection in Commonsense ReasoningChengfeng Zhao, Shizhu He, Shanshan Jiang, Bin Dong 等EMNLP 2025
它引用的顶会 Paper24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
相关 Paper
- KILM: Knowledge Injection into Encoder-Decoder Language ModelsYan Xu, Mahdi Namazifar, Devamanyu Hazarika, Aishwarya Padmakumar 等ACL 2023 · 被引用 16 次
- Knowledge Prompting in Pre-trained Language Model for Natural Language UnderstandingJianing Wang, Wenkang Huang, Minghui Qiu, Qiuhui Shi 等EMNLP 2022 · 被引用 26 次
- Plug-and-Play Knowledge Injection for Pre-trained Language ModelsZhengyan Zhang, Zhiyuan Zeng, Yankai Lin, Huadong Wang 等ACL 2023 · 被引用 10 次
- Thrust: Adaptively Propels Large Language Models with External KnowledgeXinran Zhao, Hongming Zhang, Xiaoman Pan, Wenlin Yao 等NeurIPS 2023 · 被引用 5 次
- Pre-training Language Models with Deterministic Factual KnowledgeShaobo Li, Xiaoguang Li, Lifeng Shang, Chengjie Sun 等EMNLP 2022 · 被引用 13 次
