KILM: Knowledge Injection into Encoder-Decoder Language Models
Yan Xu, Mahdi Namazifar, Devamanyu Hazarika, Aishwarya Padmakumar, Yang Liu, Dilek Hakkani-Tür
Abstract
Large pre-trained language models (PLMs) have been shown to retain implicit knowledge within their parameters. To enhance this implicit knowledge, we propose Knowledge Injection into Language Models (KILM), a novel approach that injects entity-related knowledge into encoder-decoder PLMs, via a generative knowledge infilling objective through continued pre-training. This is done without architectural modifications to the PLMs or adding additional parameters. Experimental results over a suite of knowledge-intensive tasks spanning numerous datasets show that KILM enables models to retain more knowledge and hallucinate less while preserving their original performance on general NLU and NLG tasks. KILM also demonstrates improved zero-shot performances on tasks such as entity disambiguation, outperforming state-of-the-art models having 30x more parameters. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a67da05b-82cd-4492-bdfc-c6e5d7f0d601Cited by top-tier papers7
- Propagating Knowledge Updates to LMs Through DistillationShankar Padmanabhan, Yasumasa Onoe, Michael J. Q. Zhang, Greg Durrett et al.NeurIPS 2023 · 33 citations
- EvolveBench: A Comprehensive Benchmark for Assessing Temporal Awareness in LLMs on Evolving KnowledgeZhiyuan Zhu, Yusheng Liao, Zhe Chen, Yuhao Wang et al.ACL 2025 · 10 citations
- Synthetic Knowledge Ingestion: Towards Knowledge Refinement and Injection for Enhancing Large Language ModelsJiaxin Zhang, Wendi Cui, Yiran Huang, Kamalika Das et al.EMNLP 2024 · 8 citations
- Structure-aware Domain Knowledge Injection for Large Language ModelsKai Liu, Ze Chen, Zhihang Fu, Wei Zhang et al.ACL 2025 · 5 citations
- Increasing Coverage and Precision of Textual Information in Multilingual Knowledge GraphsSimone Conia, Min Li, Daniel Lee, Umar Farooq Minhas et al.EMNLP 2023 · 3 citations
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- ERNIE 2.0: A Continual Pre-Training Framework for Language UnderstandingYu Sun, Shuohuan Wang, Yu-Kun Li, Shikun Feng et al.AAAI 2020 · 885 citations
- LUKE: Deep Contextualized Entity Representations with Entity-aware Self-attentionIkuya Yamada, Akari Asai, Hiroyuki Shindo, Hideaki Takeda et al.EMNLP 2020 · 562 citations
- Scalable Zero-shot Entity Linking with Dense Entity RetrievalLedell Wu, Fabio Petroni, Martin Josifoski, Sebastian Riedel et al.EMNLP 2020 · 336 citations
Related papers
- DKPLM: Decomposable Knowledge-Enhanced Pre-trained Language Model for Natural Language UnderstandingTaolin Zhang, Chengyu Wang, Nan Hu, Minghui Qiu et al.AAAI 2022 · 36 citations
- Knowledge Rumination for Pre-trained Language ModelsYunzhi Yao, Peng Wang, Shengyu Mao, Chuanqi Tan et al.EMNLP 2023 · 2 citations
- Enhancing Multilingual Language Model with Massive Multilingual Knowledge TriplesLinlin Liu, Xin Li, Ruidan He, Lidong Bing et al.EMNLP 2022 · 15 citations
- XLM-K: Improving Cross-Lingual Language Model Pre-training with Multilingual KnowledgeXiaoze Jiang, Yaobo Liang, Weizhu Chen, Nan DuanAAAI 2022 · 31 citations
- A Unified Encoder-Decoder Framework with Entity MemoryZhihan Zhang, Wenhao Yu, Chenguang Zhu, Meng JiangEMNLP 2022 · 9 citations
