HET-KG: Communication-Efficient Knowledge Graph Embedding Training via Hotness-Aware Cache
Sicong Dong, Xupeng Miao, Pengkai Liu, Xin Wang, Bin Cui, Jianxin Li
Abstract
With the popularization and application of Artificial Intelligence technology, knowledge graph embedding methods are widely used for a variety of machine learning tasks. However, most of the current knowledge graph embedding models are trained with a large number of parameters and high computational time complexity. This becomes a main obstacle to apply these existing models to large-scale knowledge graphs. To address this challenge, we propose HET-KG, a distributed system for training knowledge graph embedding efficiently. HET-KG can reduce the communication overheads by introducing a cache embedding table structure to maintain hot-embeddings at each worker. To improve the effectiveness of the cache mechanism, we design a prefetching algorithm and a filtering algorithm for adaptively selecting hot-embeddings, and provide two kinds of hot-embedding table construction strategies. To address the issue of inconsistency between the local cached hot-embeddings and the global embeddings, we also develop a hot-embedding synchronization algorithm for dynamically updating the cache embedding table, which can guarantee the inconsistency bounded within a given threshold. Finally, extensive experiments are conducted on three knowledge graph datasets FB15k, WN18, and Freebase-86m. The experimental results show that HET-KG achieves 3.7x and 1.1x speedup over the state-of-the-art systems PyTorch-BigGraph and DGL-KE, respectively.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers5
- Distributed Graph Embedding with Information-Oriented Random WalksPeng Fang, Arijit Khan, Siqiang Luo, Fang Wang et al.VLDB 2023 · 18 citations
- HySAE: An Efficient Semantic-Enhanced Representation Learning Model for Knowledge Hypergraph Link PredictionZhao Li, Xin Wang, Jun Zhao, Feng Feng et al.WWW 2025 · 14 citations
- TIGER: Training Inductive Graph Neural Network for Large-scale Knowledge Graph ReasoningKai Wang, Yuwei Xu, Siqiang LuoVLDB 2024 · 3 citations
- Scalable Feature Learning on Huge Knowledge Graphs for Downstream Machine LearningFélix Lefebvre, Gaël VaroquauxNeurIPS 2025 · 1 citation
- OMeGa: Boosting Large-scale Graph Embeddings with Heterogeneous Memory ProcessingPeng Fang, Siqiang Luo, Fang Wang, Bolong Zheng et al.ICDE 2025 · 1 citation
Related papers
- DGL-KE: Training Knowledge Graph Embeddings at ScaleDa Zheng, Xiang Song, Chao Ma, Zeyuan Tan et al.SIGIR 2020 · 132 citations
- Parallel Training of Knowledge Graph Embedding Models: A Comparison of TechniquesAdrian Kochsiek, Rainer GemullaVLDB 2022 · 33 citations
- HET-GMP: A Graph-based System Approach to Scaling Large Embedding Model TrainingXupeng Miao, Yining Shi, Hailin Zhang, Xin Zhang et al.SIGMOD 2022 · 24 citations
- SIT: Selective Incremental Training for Dynamic Knowledge Graph EmbeddingZhifeng Jia, Hanmo Liu, Haoyang Li, Lei ChenICDE 2025 · 1 citation
- HET: Scaling out Huge Embedding Model Training via Cache-enabled Distributed FrameworkXupeng Miao, Hailin Zhang, Yining Shi, Xiaonan Nie et al.VLDB 2022 · 70 citations
