OntoTune: Ontology-Driven Self-training for Aligning Large Language Models
Zhiqiang Liu, Chengtao Gan, Junjie Wang, Yichi Zhang, Zhongpu Bo, Mengshu Sun, Huajun Chen, Wen Zhang
Abstract
Existing domain-specific Large Language Models (LLMs) are typically developed by fine-tuning general-purposed LLMs with large-scale domain-specific corpora. However, training on large-scale corpora often fails to effectively organize domain knowledge of LLMs, leading to fragmented understanding. Inspired by how humans connect concepts and organize knowledge through mind maps, we aim to emulate this approach by using ontology with hierarchical conceptual knowledge to reorganize LLM's domain knowledge. From this perspective, we propose an ontology-driven self-training framework called OntoTune, which aims to align LLMs with ontology through in-context learning, enabling the generation of responses guided by the ontology. We leverage in-context learning to identify whether the LLM has acquired the specific concept's ontology knowledge, and select the entries not yet mastered by LLM as the training set to further align the LLM with ontology. Compared to existing domain LLMs based on newly collected large-scale domain-specific corpora, our OntoTune, which relies on the existing, long-term developed ontology and LLM itself, significantly reduces data maintenance costs and offers improved generalization ability. We conduct our study in the medical domain to evaluate the effectiveness of OntoTune, utilizing a standardized medical ontology, SNOMED CT as our ontology source. Experimental results demonstrate that OntoTune achieves state-of-the-art performance in both in-ontology task hypernym discovery and out-of-ontology task medical domain QA. Moreover, compared to the latest direct ontology injection method TaxoLLaMA, our OntoTune better preserves original knowledge of LLM. The code and data are available at https://github.com/zjukg/OntoTune.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2c1c1f2c-6af3-4838-a915-3faec841a49bCited by top-tier papers4
- UniHR: Hierarchical Representation Learning for Unified Knowledge Graph Link PredictionZhiqiang Liu, Yin Hua, Mingyang Chen, Yichi Zhang et al.AAAI 2026 · 5 citations
- JUDO: A Juxtaposed Domain-Oriented Multimodal Reasoner for Industrial Anomaly QAHyunju Kang, Woohyun Lee, Jaewon Kim, Hogun ParkICLR 2026 · 4 citations
- Collaboration of Fusion and Independence: Hypercomplex-driven Robust Multi-Modal Knowledge Graph CompletionZhiqiang Liu, Yichi Zhang, Mengshu Sun, Lei Liang et al.ACL 2026 · 1 citation
- Imprint of the Forgotten: Stealthy Membership Inference in Unlearned Graph Neural NetworksHe Zhang, Bang Wu, Xiaoning Liu, Karin Verspoor et al.AAAI 2026
Builds on16
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- LIMA: Less Is More for AlignmentChunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer et al.NeurIPS 2023 · 1,486 citations
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 1,126 citations
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen et al.ICLR 2024 · 1,104 citations
Related papers
- Structure-aware Domain Knowledge Injection for Large Language ModelsKai Liu, Ze Chen, Zhihang Fu, Wei Zhang et al.ACL 2025 · 5 citations
- The Missing Alignment Link of In-context Learning on SequencesHarshvardhan Agarwal, Sunita SarawagiICML 2025
- Investigating and Mitigating Catastrophic Forgetting in Medical Knowledge Injection through Internal Knowledge Augmentation LearningYuxuan Zhou, Xien Liu, Xiao Zhang, Chen Ning et al.NeurIPS 2025 · 5 citations
- SciCustom: A Framework for Custom Evaluation of Scientific Capabilities in Large Language ModelsYiyang Gu, Junwei Yang, Junyu Luo, Ye Yuan et al.ACL 2026
- End-to-End Ontology Learning with Large Language ModelsAndy Lo, Albert Q. Jiang, Wenda Li, Mateja JamnikNeurIPS 2024 · 33 citations
