HET-GMP: A Graph-based System Approach to Scaling Large Embedding Model Training
Xupeng Miao, Yining Shi, Hailin Zhang, Xin Zhang, Xiaonan Nie, Zhi Yang, Bin Cui
摘要
Embedding models have been recognized as an effective learning paradigm for high-dimensional data. However, a major embedding model training obstacle is that updating and retrieving the shared large-scale embedding parameters usually dominates the distributed training cycle, leading to significant scalability issues. This paper presents HET-GMP, a distributed system on training embedding models. Uniquely, HET-GMP takes advantage of a graph-based approach to efficiently increase scalability. The key insight guiding our design is the "graph way of thinking". HET-GMP creates a bigraph abstraction to represent the access relationships between data samples and embedding vectors. This enables HET-GMP to embrace graph locality and skewness as new performance opportunities and to exploit graph-based replication/partitioning and bounded-asynchronous synchronization to reduce communication overhead. We evaluate the system on the embedding models for click-through rate (CTR) prediction, which presents the most significant challenge and communication bottleneck due to heavy access concurrency to a huge embedding table. The result shows that HET-GMP supports embedding model training with 10 11 parameters, achieving a reduction in communication up to 87.5% and an up-to 27.5× speedup over the state-of-the-art baseline systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Galvatron: Efficient Transformer Training over Multiple GPUs Using Automatic ParallelismXupeng Miao, Yujie Wang, Youhe Jiang, Chunan Shi 等VLDB 2023 · 被引用 113 次
- FlexMoE: Scaling Large-scale Sparse Pre-trained Model Training via Dynamic Device PlacementXiaonan Nie, Xupeng Miao, Zilong Wang, Zichao Yang 等SIGMOD 2023 · 被引用 40 次
- Accelerating Sampling and Aggregation Operations in GNN Frameworks with GPU Initiated Direct Storage AccessesJeongmin Brian Park, Vikram Sharma Mailthody, Zaid Qureshi, Wen-Mei HwuVLDB 2024 · 被引用 37 次
- ETC: Efficient Training of Temporal Graph Neural Networks over Large-scale Dynamic GraphsShihong Gao, Yiming Li, Yanyan Shen, Yingxia Shao 等VLDB 2024 · 被引用 32 次
- Towards Communication-efficient Vertical Federated Learning Training via Cache-enabled Local UpdateFangcheng Fu, Xupeng Miao, Jiawei Jiang, Huanran Xue 等VLDB 2022 · 被引用 31 次
它引用的顶会 Paper9
- DGL-KE: Training Knowledge Graph Embeddings at ScaleDa Zheng, Xiang Song, Chao Ma, Zeyuan Tan 等SIGIR 2020 · 被引用 132 次
- Reliable Data Distillation on Graph Convolutional NetworkWentao Zhang, Xupeng Miao, Yingxia Shao, Jiawei Jiang 等SIGMOD 2020 · 被引用 72 次
- HET: Scaling out Huge Embedding Model Training via Cache-enabled Distributed FrameworkXupeng Miao, Hailin Zhang, Yining Shi, Xiaonan Nie 等VLDB 2022 · 被引用 70 次
- Heterogeneity-Aware Distributed Machine Learning Training via Partial ReduceXupeng Miao, Xiaonan Nie, Yingxia Shao, Zhi Yang 等SIGMOD 2021 · 被引用 64 次
- Agile and Accurate CTR Prediction Model Training for Massive-Scale Online Advertising SystemsZhiqiang Xu, Dong Li, Weijie Zhao, Xing Shen 等SIGMOD 2021 · 被引用 38 次
相关 Paper
- ScaleFreeCTR: MixCache-based Distributed Training System for CTR Models with Huge Embedding TableHuifeng Guo, Wei Guo, Yong Gao, Ruiming Tang 等SIGIR 2021 · 被引用 24 次
- HET-KG: Communication-Efficient Knowledge Graph Embedding Training via Hotness-Aware CacheSicong Dong, Xupeng Miao, Pengkai Liu, Xin Wang 等ICDE 2022 · 被引用 13 次
- FeLoG: Scalable and Efficient Distributed Graph Embedding with Feedback Loop MechanismPeng Fang, Arijit Khan, Ziqiang Wu, Zhenli Li 等VLDB 2026
- Distributed Graph Embedding with Information-Oriented Random WalksPeng Fang, Arijit Khan, Siqiang Luo, Fang Wang 等VLDB 2023 · 被引用 18 次
- HypeReca: Distributed Heterogeneous In-Memory Embedding Database for Training Recommender ModelsJiaao He, Shengqi Chen, Kezhao Huang, Jidong ZhaiUSENIX ATC 2025 · 被引用 2 次
