Parallel Training of Knowledge Graph Embedding Models: A Comparison of Techniques
Adrian Kochsiek, Rainer Gemulla
Abstract
Knowledge graph embedding (KGE) models represent the entities and relations of a knowledge graph (KG) using dense continuous representations called embeddings. KGE methods have recently gained traction for tasks such as knowledge graph completion and reasoning as well as to provide suitable entity representations for downstream learning tasks. While a large part of the available literature focuses on small KGs, a number of frameworks that are able to train KGE models for large-scale KGs by parallelization across multiple GPUs or machines have recently been proposed. So far, the benefits and drawbacks of the various parallelization techniques have not been studied comprehensively. In this paper, we report on an experimental study in which we presented, re-implemented in a common computational framework, investigated, and improved the available techniques. We found that the evaluation methodologies used in prior work are often not comparable and can be misleading, and that most of currently implemented training methods tend to have a negative impact on embedding quality. We propose a simple but effective variation of the stratification technique used by PyTorch BigGraph for mitigation. Moreover, basic random partitioning can be an effective or even the best-performing choice when combined with suitable sampling techniques. Ultimately, we found that efficient and effective parallel training of large-scale KGE models is indeed achievable but requires a careful choice of techniques.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cd4e8594-056b-4e81-a4c9-21c22612c76fCited by top-tier papers7
- Sequence-to-Sequence Knowledge Graph Completion and Question AnsweringApoorv Saxena, Adrian Kochsiek, Rainer GemullaACL 2022 · 183 citations
- Experimental Analysis of Large-scale Learnable Vector Storage CompressionHailin Zhang, Penghao Zhao, Xupeng Miao, Yingxia Shao et al.VLDB 2024 · 20 citations
- NuPS: A Parameter Server for Machine Learning with Non-Uniform Parameter AccessAlexander Renz-Wieland, Rainer Gemulla, Zoi Kaoudi, Volker MarklSIGMOD 2022 · 19 citations
- Scalable Graph Convolutional Network Training on Distributed-Memory SystemsGunduz Vehbi Demirci, Aparajita Haldar, Hakan FerhatosmanogluVLDB 2023 · 18 citations
- CAFE: Towards Compact, Adaptive, and Fast Embedding for Large-scale Recommendation ModelsHailin Zhang, Zirui Liu, Boxuan Chen, Yikai Zhao et al.SIGMOD 2024 · 15 citations
Builds on4
- You CAN Teach an Old Dog New Tricks! On Training Knowledge Graph EmbeddingsDaniel Ruffinelli, Samuel Broscheit, Rainer GemullaICLR 2020 · 238 citations
- DGL-KE: Training Knowledge Graph Embeddings at ScaleDa Zheng, Xiang Song, Chao Ma, Zeyuan Tan et al.SIGIR 2020 · 132 citations
- HittER: Hierarchical Transformers for Knowledge Graph EmbeddingsSanxing Chen, Xiaodong Liu, Jianfeng Gao, Jian Jiao et al.EMNLP 2021 · 110 citations
- Dynamic Parameter Allocation in Parameter ServersAlexander Renz-Wieland, Rainer Gemulla, Steffen Zeuch, Volker MarklVLDB 2020 · 18 citations
Related papers
- HET-KG: Communication-Efficient Knowledge Graph Embedding Training via Hotness-Aware CacheSicong Dong, Xupeng Miao, Pengkai Liu, Xin Wang et al.ICDE 2022 · 13 citations
- KGDM: A Diffusion Model to Capture Multiple Relation Semantics for Knowledge Graph EmbeddingXiao Long, Liansheng Zhuang, Aodi Li, Jiuchang Wei et al.AAAI 2024 · 16 citations
- Distributed Graph Embedding with Information-Oriented Random WalksPeng Fang, Arijit Khan, Siqiang Luo, Fang Wang et al.VLDB 2023 · 18 citations
- Efficient Non-Sampling Knowledge Graph EmbeddingZelong Li, Jianchao Ji, Zuohui Fu, Yingqiang Ge et al.WWW 2021 · 41 citations
- SAGE: Scale-Aware Gradual Evolution for Continual Knowledge Graph EmbeddingYifei Li, Lingling Zhang, Hang Yan, Tianzhe Zhao et al.KDD 2025 · 2 citations
