HET-GMP: A Graph-based System Approach to Scaling Large Embedding Model Training
Xupeng Miao, Yining Shi, Hailin Zhang, Xin Zhang, Xiaonan Nie, Zhi Yang, Bin Cui
Abstract
Embedding models have been recognized as an effective learning paradigm for high-dimensional data. However, a major embedding model training obstacle is that updating and retrieving the shared large-scale embedding parameters usually dominates the distributed training cycle, leading to significant scalability issues. This paper presents HET-GMP, a distributed system on training embedding models. Uniquely, HET-GMP takes advantage of a graph-based approach to efficiently increase scalability. The key insight guiding our design is the "graph way of thinking". HET-GMP creates a bigraph abstraction to represent the access relationships between data samples and embedding vectors. This enables HET-GMP to embrace graph locality and skewness as new performance opportunities and to exploit graph-based replication/partitioning and bounded-asynchronous synchronization to reduce communication overhead. We evaluate the system on the embedding models for click-through rate (CTR) prediction, which presents the most significant challenge and communication bottleneck due to heavy access concurrency to a huge embedding table. The result shows that HET-GMP supports embedding model training with 10 11 parameters, achieving a reduction in communication up to 87.5% and an up-to 27.5× speedup over the state-of-the-art baseline systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9bf4813d-8d7f-4c7f-908e-21fba91bf7acCited by top-tier papers16
- Galvatron: Efficient Transformer Training over Multiple GPUs Using Automatic ParallelismXupeng Miao, Yujie Wang, Youhe Jiang, Chunan Shi et al.VLDB 2023 · 113 citations
- FlexMoE: Scaling Large-scale Sparse Pre-trained Model Training via Dynamic Device PlacementXiaonan Nie, Xupeng Miao, Zilong Wang, Zichao Yang et al.SIGMOD 2023 · 40 citations
- Accelerating Sampling and Aggregation Operations in GNN Frameworks with GPU Initiated Direct Storage AccessesJeongmin Brian Park, Vikram Sharma Mailthody, Zaid Qureshi, Wen-Mei HwuVLDB 2024 · 37 citations
- ETC: Efficient Training of Temporal Graph Neural Networks over Large-scale Dynamic GraphsShihong Gao, Yiming Li, Yanyan Shen, Yingxia Shao et al.VLDB 2024 · 32 citations
- Towards Communication-efficient Vertical Federated Learning Training via Cache-enabled Local UpdateFangcheng Fu, Xupeng Miao, Jiawei Jiang, Huanran Xue et al.VLDB 2022 · 31 citations
Builds on9
- DGL-KE: Training Knowledge Graph Embeddings at ScaleDa Zheng, Xiang Song, Chao Ma, Zeyuan Tan et al.SIGIR 2020 · 132 citations
- Reliable Data Distillation on Graph Convolutional NetworkWentao Zhang, Xupeng Miao, Yingxia Shao, Jiawei Jiang et al.SIGMOD 2020 · 72 citations
- HET: Scaling out Huge Embedding Model Training via Cache-enabled Distributed FrameworkXupeng Miao, Hailin Zhang, Yining Shi, Xiaonan Nie et al.VLDB 2022 · 70 citations
- Heterogeneity-Aware Distributed Machine Learning Training via Partial ReduceXupeng Miao, Xiaonan Nie, Yingxia Shao, Zhi Yang et al.SIGMOD 2021 · 64 citations
- Agile and Accurate CTR Prediction Model Training for Massive-Scale Online Advertising SystemsZhiqiang Xu, Dong Li, Weijie Zhao, Xing Shen et al.SIGMOD 2021 · 38 citations
Related papers
- ScaleFreeCTR: MixCache-based Distributed Training System for CTR Models with Huge Embedding TableHuifeng Guo, Wei Guo, Yong Gao, Ruiming Tang et al.SIGIR 2021 · 24 citations
- HET-KG: Communication-Efficient Knowledge Graph Embedding Training via Hotness-Aware CacheSicong Dong, Xupeng Miao, Pengkai Liu, Xin Wang et al.ICDE 2022 · 13 citations
- FeLoG: Scalable and Efficient Distributed Graph Embedding with Feedback Loop MechanismPeng Fang, Arijit Khan, Ziqiang Wu, Zhenli Li et al.VLDB 2026
- Distributed Graph Embedding with Information-Oriented Random WalksPeng Fang, Arijit Khan, Siqiang Luo, Fang Wang et al.VLDB 2023 · 18 citations
- HypeReca: Distributed Heterogeneous In-Memory Embedding Database for Training Recommender ModelsJiaao He, Shengqi Chen, Kezhao Huang, Jidong ZhaiUSENIX ATC 2025 · 2 citations
