Progressive Distillation Based on Masked Generation Feature Method for Knowledge Graph Completion
Cunhang Fan, Yujie Chen, Jun Xue, Yonghui Kong, Jianhua Tao, Zhao Lv
Abstract
In recent years, knowledge graph completion (KGC) models based on pre-trained language model (PLM) have shown promising results. However, the large number of parameters and high computational cost of PLM models pose challenges for their application in downstream tasks. This paper proposes a progressive distillation method based on masked generation features for KGC task, aiming to significantly reduce the complexity of pre-trained models. Specifically, we perform pre-distillation on PLM to obtain high-quality teacher models, and compress the PLM network to obtain multi-grade student models. However, traditional feature distillation suffers from the limitation of having a single representation of information in teacher models. To solve this problem, we propose masked generation of teacher-student features, which contain richer representation information. Furthermore, there is a significant gap in representation ability between teacher and student. Therefore, we design a progressive distillation method to distill student models at each grade level, enabling efficient knowledge transfer from teachers to students. The experimental results demonstrate that the model in the pre-distillation stage surpasses the existing state-of-the-art methods. Furthermore, in the progressive distillation stage, the model significantly reduces the model parameters while maintaining a certain level of performance. Specifically, the model parameters of the lower-grade student model are reduced by 56.7% compared to the baseline.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0d0c8d32-6baf-49e5-8b28-4c364e6dbdaeCited by top-tier papers3
- Progressive Mask Distillation for Self-supervised Video RepresentationKewei Wu, Chong Liang, Zhao Xie, Dan GuoCVPR 2026
- Dark Side of Modalities: Reinforced Multimodal Distillation for Multimodal Knowledge Graph ReasoningYu Zhao, Ying Zhang, Xuhui Sui, Baohang Zhou et al.ACM MM 2025
- Croppable Knowledge Graph EmbeddingYushan Zhu, Wen Zhang, Zhiqiang Liu, Mingyang Chen et al.ACL 2025
Builds on7
- Composition-based Multi-Relational Graph Convolutional NetworksShikhar Vashishth, Soumya Sanyal, Vikram Nitin, Partha P. TalukdarICLR 2020 · 1,105 citations
- Decoupled Knowledge DistillationBorui Zhao, Quan Cui, Renjie Song, Yiyu Qiu et al.CVPR 2022 · 835 citations
- Improving Multi-hop Question Answering over Knowledge Graphs using Knowledge Base EmbeddingsApoorv Saxena, Aditay Tripathi, Partha P. TalukdarACL 2020 · 488 citations
- Focal and Global Knowledge Distillation for DetectorsZhendong Yang, Zhe Li, Xiaohu Jiang, Yuan Gong et al.CVPR 2022 · 325 citations
- MulDE: Multi-teacher Knowledge Distillation for Low-dimensional Knowledge Graph EmbeddingsKai Wang, Yu Liu, Qian Ma, Quan Z. ShengWWW 2021 · 67 citations
Related papers
- Meta-KD: A Meta Knowledge Distillation Framework for Language Model Compression across DomainsHaojie Pan, Chengyu Wang, Minghui Qiu, Yichang Zhang et al.ACL 2021
- Joint Pre-Encoding Representation and Structure Embedding for Efficient and Low-Resource Knowledge Graph CompletionChenyu Qiu, Pengjiang Qian, Chuang Wang, Jian Yao et al.EMNLP 2024 · 2 citations
- RaSE-KGC: A Relation-Aware Segment Encoding Approach for Knowledge Graph CompletionChenxiao Lin, Ye Luo, Kunhong Liu, Qingqiang WuICDE 2026
- Towards Efficient Pre-Trained Language Model via Feature Correlation DistillationKun Huang, Xin Guo, Meng WangNeurIPS 2023 · 8 citations
- Joint Pre-training and Local Re-training: Transferable Representation Learning on Multi-source Knowledge GraphsZequn Sun, Jiacheng Huang, Jinghao Lin, Xiaozhou Xu et al.KDD 2023 · 5 citations
