Synthetic Knowledge Ingestion: Towards Knowledge Refinement and Injection for Enhancing Large Language Models
Jiaxin Zhang, Wendi Cui, Yiran Huang, Kamalika Das, Kumar Sricharan
摘要
Large language models (LLMs) are proficient in capturing factual knowledge across various domains. However, refining their capabilities on previously seen knowledge or integrating new knowledge from external sources remains a significant challenge. In this work, we propose a novel synthetic knowledge ingestion method called Ski, which leverages fine-grained synthesis, interleaved generation, and assemble augmentation strategies to construct high-quality data representations from raw knowledge sources. We then integrate Ski and its variations with three knowledge injection techniques: Retrieval Augmented Generation (RAG), Supervised Fine-tuning (SFT), and Continual Pre-training (CPT) to inject and refine knowledge in language models. Extensive empirical experiments are conducted on various question-answering tasks spanning finance, biomedicine, and open-generation domains to demonstrate that Ski significantly outperforms baseline methods by facilitating effective knowledge injection. We believe that our work is an important step towards enhancing the factual accuracy of LLM outputs by refining knowledge representation and injection capabilities. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Knowledge Boundary of Large Language Models: A SurveyMoxin Li, Yong Zhao, Wenxuan Zhang, Shuaiyi Li 等ACL 2025 · 被引用 33 次
- Rote Learning Considered Useful: Generalizing over Memorized Data in LLMsQinyuan Wu, Soumi Das, Mahsa Amani, Bishwamittra Ghosh 等ICLR 2026 · 被引用 6 次
它引用的顶会 Paper14
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- LIMA: Less Is More for AlignmentChunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer 等NeurIPS 2023 · 被引用 1,486 次
- Large Language Models Struggle to Learn Long-Tail KnowledgeNikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace 等ICML 2023 · 被引用 623 次
相关 Paper
- Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMsOded Ovadia, Menachem Brief, Moshik Mishaeli, Oren ElishaEMNLP 2024 · 被引用 89 次
- Embedding Domain Knowledge for Large Language Models via Reinforcement Learning from Augmented GenerationChaojun Nie, Jun Zhou, Guanxiang Wang, Shisong Wu 等EMNLP 2025
- SR-KI: Scalable and Real-Time Knowledge Integration into LLMs via Supervised AttentionBohan Yu, Wei Huang, Kang LiuAAAI 2026
- Propagating Knowledge Updates to LMs Through DistillationShankar Padmanabhan, Yasumasa Onoe, Michael J. Q. Zhang, Greg Durrett 等NeurIPS 2023 · 被引用 33 次
- In-depth Analysis of Graph-based RAG in a Unified FrameworkYingli Zhou, Yaodong Su, Youran Sun, Shu Wang 等VLDB 2025 · 被引用 48 次
