Structure-aware Domain Knowledge Injection for Large Language Models
Kai Liu, Ze Chen, Zhihang Fu, Wei Zhang, Rongxin Jiang, Fan Zhou, Yaowu Chen, Yue Wu, Jieping Ye
摘要
This paper introduces a pioneering methodology, termed StructTuning, to efficiently transform foundation Large Language Models (LLMs) into domain specialists. It significantly reduces the training corpus needs to a mere 5% while achieving an impressive 100% of traditional knowledge injection performance. Motivated by structured human education, we propose a novel two-stage strategy for knowledge injection and alignment: Structure-aware Continual Pre-Training (SCPT) and Structureaware Supervised Fine-Tuning (SSFT). In the SCPT phase, we automatically extract the domain knowledge taxonomy and reorganize the training corpora, enabling LLMs to effectively link textual segments to targeted knowledge points within the taxonomy. In the SSFT phase, we explicitly prompt models to elucidate the underlying knowledge structure in their outputs, leveraging the structured domain insight to address practical problems. Our ultimate method was extensively evaluated across model architectures and scales on LongBench and MMed-Bench datasets, demonstrating superior performance against other knowledge injection methods. We also explored our method's scalability across different training corpus sizes, laying the foundation to enhance domain-specific LLMs with better data utilization. Code
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Knowledge Boundary of Large Language Models: A SurveyMoxin Li, Yong Zhao, Wenxuan Zhang, Shuaiyi Li 等ACL 2025 · 被引用 33 次
- Collaboration of Large Language Models and Small Recommendation Models for Device-Cloud RecommendationZheqi Lv, Tianyu Zhan, Wenjie Wang, Xinyu Lin 等KDD 2025 · 被引用 4 次
- Structural Entropy Guided Agent for Detecting and Repairing Knowledge Deficiencies in LLMsYifan Wei, Xiaoyan Yu, Tengfei Pan, Angsheng Li 等NeurIPS 2025 · 被引用 3 次
- Analyzing and Internalizing Complex Policy Documents for LLM AgentsJiateng Liu, Zhenhailong Wang, Xiaojiang Huang, Yingjie Li 等ACL 2026
- Improving Long-Context Summarization with Multi-Granularity Retrieval OptimizationXueyu Chen, Kaitao Song, Zifan Song, Dongsheng Li 等AAAI 2026
它引用的顶会 Paper16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- K-BERT: Enabling Language Representation with Knowledge GraphWeijie Liu, Peng Zhou, Zhe Zhao, Zhiruo Wang 等AAAI 2020 · 被引用 898 次
相关 Paper
- SciLitLLM: How to Adapt LLMs for Scientific Literature UnderstandingSihang Li, Jin Huang, Jiaxi Zhuang, Yaorui Shi 等ICLR 2025 · 被引用 1 次
- Mix-CPT: A Domain Adaptation Framework via Decoupling Knowledge Learning and Format AlignmentJinhao Jiang, Junyi Li, Xin Zhao, Yang Song 等ICLR 2025
- OntoTune: Ontology-Driven Self-training for Aligning Large Language ModelsZhiqiang Liu, Chengtao Gan, Junjie Wang, Yichi Zhang 等WWW 2025 · 被引用 11 次
- Learning or Self-aligning? Rethinking Instruction Fine-tuningMengjie Ren, Boxi Cao, Hongyu Lin, Cao Liu 等ACL 2024
- ADEPT: Continual Pretraining via Adaptive Expansion and Dynamic Decoupled TuningJinyang Zhang, Yue Fang, Hongxin Ding, Weibin Liao 等ICLR 2026 · 被引用 5 次
