How Training Data Shapes the Use of Parametric and In-Context Knowledge in Language Models
Minsung Kim, Dong-Kyum Kim, Jea Kwon, Nakyeong Yang, Kyomin Jung, Meeyoung Cha
摘要
Large language models leverage both parametric knowledge acquired during pretraining and in-context knowledge provided at inference time. Crucially, when these sources conflict, models arbitrate based on their internal confidence, preferring parametric knowledge for high-confidence facts while deferring to context for less familiar ones. However, the training conditions that give rise to these fundamental behaviors remain unclear. Here we conduct controlled experiments using synthetic corpora to identify the specific data properties that shape knowledge utilization. Our results reveal a counterintuitive finding: the robust, balanced use of both knowledge sources is an emergent property that requires the cooccurrence of three factors typically considered detrimental, including (i) intra-document repetition, (ii) a moderate degree of intra-document inconsistency, and (iii) a skewed knowledge distribution. We further show that these dynamics arise in real-world language model pretraining and analyze how post-training procedures reshape arbitration strategies. Together, our findings provide empirical guidance for designing training data that supports the reliable integration of parametric and in-context knowledge in language models. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- Data Distributional Properties Drive Emergent In-Context Learning in TransformersStephanie C. Y. Chan, Adam Santoro, Andrew K. Lampinen, Jane X. Wang 等NeurIPS 2022 · 被引用 407 次
- Adaptive Chameleon or Stubborn Sloth: Revealing the Behavior of Large Language Models in Knowledge ConflictsJian Xie, Kai Zhang, Jiangjie Chen, Renze Lou 等ICLR 2024 · 被引用 294 次
相关 Paper
- Formality is Favored: Unraveling the Learning Preferences of Large Language Models on Data with Conflicting KnowledgeJiahuan Li, Yiqing Cao, Shujian Huang, Jiajun ChenEMNLP 2024
- When Context Leads but Parametric Memory Follows in Large Language ModelsYufei Tao, Adam Hiatt, Erik Haake, Antonie J. Jetter 等EMNLP 2024 · 被引用 4 次
- Entity-Based Knowledge Conflicts in Question AnsweringShayne Longpre, Kartik Perisetla, Anthony Chen, Nikhil Ramesh 等EMNLP 2021 · 被引用 3 次
- Whose Facts Win? LLM Source Preferences under Knowledge ConflictsJakob Schuster, Vagrant Gautam, Katja MarkertACL 2026 · 被引用 3 次
- Rich Knowledge Sources Bring Complex Knowledge Conflicts: Recalibrating Models to Reflect Conflicting EvidenceHung-Ting Chen, Michael J. Q. Zhang, Eunsol ChoiEMNLP 2022 · 被引用 27 次
