MulDE: Multi-teacher Knowledge Distillation for Low-dimensional Knowledge Graph Embeddings
Kai Wang, Yu Liu, Qian Ma, Quan Z. Sheng
Abstract
Link prediction based on knowledge graph embeddings (KGE) aims to predict new triples to automatically construct knowledge graphs (KGs). However, recent KGE models achieve performance improvements by excessively increasing the embedding dimensions, which may cause enormous training costs and require more storage space. In this paper, instead of training high-dimensional models, we propose MulDE, a novel knowledge distillation framework, which includes multiple low-dimensional hyperbolic KGE models as teachers and two student components, namely Junior and Senior. Under a novel iterative distillation strategy, the Junior component, a low-dimensional KGE model, asks teachers actively based on its preliminary prediction results, and the Senior component integrates teachers' knowledge adaptively to train the Junior component based on two mechanisms: relation-specific scaling and contrast attention. The experimental results show that MulDE can effectively improve the performance and training speed of lowdimensional KGE models. The distilled 32-dimensional model is competitive compared to the state-of-the-art high-dimensional methods on several widely-used datasets. CCS CONCEPTS • Computing methodologies → Knowledge representation and reasoning; • Information systems → Entity relationship models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers19
- NodePiece: Compositional and Parameter-Efficient Representations of Large Knowledge GraphsMikhail Galkin, Etienne G. Denis, Jiapeng Wu, William L. HamiltonICLR 2022 · 114 citations
- Towards Continual Knowledge Graph Embedding via Incremental DistillationJiajun Liu, Wenjun Ke, Peng Wang, Ziyu Shang et al.AAAI 2024 · 52 citations
- Boosting Graph Neural Networks via Adaptive Knowledge DistillationZhichun Guo, Chunhui Zhang, Yujie Fan, Yijun Tian et al.AAAI 2023 · 48 citations
- Link Prediction with Attention Applied on Multiple Knowledge Graph Embedding ModelsCosimo Gregucci, Mojtaba Nayyeri, Daniel Hernández, Steffen StaabWWW 2023 · 37 citations
- Distillation from Heterogeneous Models for Top-K RecommendationSeongKu Kang, Wonbin Kweon, Dongha Lee, Jianxun Lian et al.WWW 2023 · 35 citations
Builds on14
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine et al.AAAI 2020 · 1,361 citations
- Composition-based Multi-Relational Graph Convolutional NetworksShikhar Vashishth, Soumya Sanyal, Vikram Nitin, Partha P. TalukdarICLR 2020 · 1,105 citations
- You CAN Teach an Old Dog New Tricks! On Training Knowledge Graph EmbeddingsDaniel Ruffinelli, Samuel Broscheit, Rainer GemullaICLR 2020 · 238 citations
- Understanding Knowledge Distillation in Non-autoregressive Machine TranslationChunting Zhou, Jiatao Gu, Graham NeubigICLR 2020 · 235 citations
- Reinforced Negative Sampling over Knowledge Graph for RecommendationXiang Wang, Yaokun Xu, Xiangnan He, Yixin Cao et al.WWW 2020 · 209 citations
Related papers
- IterDE: An Iterative Knowledge Distillation Framework for Knowledge Graph EmbeddingsJiajun Liu, Peng Wang, Ziyu Shang, Chenxiao WuAAAI 2023 · 26 citations
- Swift and Sure: Hardness-aware Contrastive Learning for Low-dimensional Knowledge Graph EmbeddingsKai Wang, Yu Liu, Quan Z. ShengWWW 2022 · 21 citations
- Geometry Awakening: Cross-Geometry Learning Exhibits Superiority over Individual StructuresYadong Sun, Xiaofeng Cao, Yu Wang, Wei Ye et al.NeurIPS 2024 · 2 citations
- Joint Pre-training and Local Re-training: Transferable Representation Learning on Multi-source Knowledge GraphsZequn Sun, Jiacheng Huang, Jinghao Lin, Xiaozhou Xu et al.KDD 2023 · 5 citations
- MARCH: Multi-Teacher Contrastive Hypergraph DistillationRongwei Xu, Zitai Qiu, Pengfei Ding, Jia Wu et al.WWW 2026
