Parallel Training of Knowledge Graph Embedding Models: A Comparison of Techniques
Adrian Kochsiek, Rainer Gemulla
摘要
Knowledge graph embedding (KGE) models represent the entities and relations of a knowledge graph (KG) using dense continuous representations called embeddings. KGE methods have recently gained traction for tasks such as knowledge graph completion and reasoning as well as to provide suitable entity representations for downstream learning tasks. While a large part of the available literature focuses on small KGs, a number of frameworks that are able to train KGE models for large-scale KGs by parallelization across multiple GPUs or machines have recently been proposed. So far, the benefits and drawbacks of the various parallelization techniques have not been studied comprehensively. In this paper, we report on an experimental study in which we presented, re-implemented in a common computational framework, investigated, and improved the available techniques. We found that the evaluation methodologies used in prior work are often not comparable and can be misleading, and that most of currently implemented training methods tend to have a negative impact on embedding quality. We propose a simple but effective variation of the stratification technique used by PyTorch BigGraph for mitigation. Moreover, basic random partitioning can be an effective or even the best-performing choice when combined with suitable sampling techniques. Ultimately, we found that efficient and effective parallel training of large-scale KGE models is indeed achievable but requires a careful choice of techniques.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Sequence-to-Sequence Knowledge Graph Completion and Question AnsweringApoorv Saxena, Adrian Kochsiek, Rainer GemullaACL 2022 · 被引用 183 次
- Experimental Analysis of Large-scale Learnable Vector Storage CompressionHailin Zhang, Penghao Zhao, Xupeng Miao, Yingxia Shao 等VLDB 2024 · 被引用 20 次
- NuPS: A Parameter Server for Machine Learning with Non-Uniform Parameter AccessAlexander Renz-Wieland, Rainer Gemulla, Zoi Kaoudi, Volker MarklSIGMOD 2022 · 被引用 19 次
- Scalable Graph Convolutional Network Training on Distributed-Memory SystemsGunduz Vehbi Demirci, Aparajita Haldar, Hakan FerhatosmanogluVLDB 2023 · 被引用 18 次
- CAFE: Towards Compact, Adaptive, and Fast Embedding for Large-scale Recommendation ModelsHailin Zhang, Zirui Liu, Boxuan Chen, Yikai Zhao 等SIGMOD 2024 · 被引用 15 次
它引用的顶会 Paper4
- You CAN Teach an Old Dog New Tricks! On Training Knowledge Graph EmbeddingsDaniel Ruffinelli, Samuel Broscheit, Rainer GemullaICLR 2020 · 被引用 238 次
- DGL-KE: Training Knowledge Graph Embeddings at ScaleDa Zheng, Xiang Song, Chao Ma, Zeyuan Tan 等SIGIR 2020 · 被引用 132 次
- HittER: Hierarchical Transformers for Knowledge Graph EmbeddingsSanxing Chen, Xiaodong Liu, Jianfeng Gao, Jian Jiao 等EMNLP 2021 · 被引用 110 次
- Dynamic Parameter Allocation in Parameter ServersAlexander Renz-Wieland, Rainer Gemulla, Steffen Zeuch, Volker MarklVLDB 2020 · 被引用 18 次
相关 Paper
- HET-KG: Communication-Efficient Knowledge Graph Embedding Training via Hotness-Aware CacheSicong Dong, Xupeng Miao, Pengkai Liu, Xin Wang 等ICDE 2022 · 被引用 13 次
- KGDM: A Diffusion Model to Capture Multiple Relation Semantics for Knowledge Graph EmbeddingXiao Long, Liansheng Zhuang, Aodi Li, Jiuchang Wei 等AAAI 2024 · 被引用 16 次
- Distributed Graph Embedding with Information-Oriented Random WalksPeng Fang, Arijit Khan, Siqiang Luo, Fang Wang 等VLDB 2023 · 被引用 18 次
- Efficient Non-Sampling Knowledge Graph EmbeddingZelong Li, Jianchao Ji, Zuohui Fu, Yingqiang Ge 等WWW 2021 · 被引用 41 次
- SAGE: Scale-Aware Gradual Evolution for Continual Knowledge Graph EmbeddingYifei Li, Lingling Zhang, Hang Yan, Tianzhe Zhao 等KDD 2025 · 被引用 2 次
