Acquisition and Application of Novel Knowledge in Large Language Models
Ziyu Shang, Jianghan Liu, Zhizhao Luo, Peng Wang, Wenjun Ke, Jiajun Liu, Zijie Xu, Guozheng Li
摘要
Recent advancements in large language models (LLMs) have demonstrated their impressive generative capabilities, primarily due to their extensive parameterization, which enables them to encode vast knowledge. However, effectively integrating new knowledge into LLMs remains a major challenge. Current research typically first constructs novel knowledge datasets and then injects this knowledge into LLMs through various techniques. However, existing methods for constructing new datasets either rely on timestamps, which lack rigor, or use simple templates for synthesis, which are simplistic and do not accurately reflect the real world. To address this issue, we propose a novel knowledge dataset construction approach that simulates biological evolution using knowledge graphs to generate synthetic entities with diverse attributes, resulting in a dataset, NovelHuman. Systematic analysis on NovelHuman reveals that the intra-sentence position of knowledge significantly affects the acquisition of knowledge. Therefore, we introduce an intra-sentence permutation to enhance knowledge acquisition. Furthermore, given that potential conflicts exist between autoregressive (AR) training objectives and permutation-based learning, we propose PermAR, a permutation-based language modeling framework for AR models. PermAR seamlessly integrates with mainstream AR architectures, endowing them with bidirectional knowledge acquisition capabilities. Extensive experiments demonstrate the superiority of PermAR, outperforming knowledge augmentation methods by 3.3%-38%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Physics of Language Models: Part 3.1, Knowledge Storage and ExtractionZeyuan Allen-Zhu, Yuanzhi LiICML 2024 · 被引用 258 次
- Skill-it! A data-driven skills framework for understanding and training language modelsMayee F. Chen, Nicholas Roberts, Kush Bhatia, Jue Wang 等NeurIPS 2023 · 被引用 143 次
- DPOT: Auto-Regressive Denoising Operator Transformer for Large-Scale PDE Pre-TrainingZhongkai Hao, Chang Su, Songming Liu, Julius Berner 等ICML 2024 · 被引用 107 次
- The Factorization Curse: Which Tokens You Predict Underlie the Reversal Curse and MoreOuail Kitouni, Niklas Nolte, Adina Williams, Michael Rabbat 等NeurIPS 2024 · 被引用 29 次
相关 Paper
- Way to Specialist: Closing Loop Between Specialized LLM and Evolving Domain Knowledge GraphYutong Zhang, Lixing Chen, Shenghong Li, Nan Cao 等KDD 2025 · 被引用 3 次
- LinkQA: Synthesizing Diverse QA from Multiple Seeds Strongly Linked by Knowledge PointsXuemiao Zhang, Can Ren, Chengying Tu, Rongxiang Weng 等ACL 2026 · 被引用 3 次
- Enhancing Multilingual Language Model with Massive Multilingual Knowledge TriplesLinlin Liu, Xin Li, Ruidan He, Lidong Bing 等EMNLP 2022 · 被引用 15 次
- MKGL: Mastery of a Three-Word LanguageLingbing Guo, Zhongpu Bo, Zhuo Chen, Yichi Zhang 等NeurIPS 2024 · 被引用 27 次
- KILM: Knowledge Injection into Encoder-Decoder Language ModelsYan Xu, Mahdi Namazifar, Devamanyu Hazarika, Aishwarya Padmakumar 等ACL 2023 · 被引用 16 次
