HET: Scaling out Huge Embedding Model Training via Cache-enabled Distributed Framework
Xupeng Miao, Hailin Zhang, Yining Shi, Xiaonan Nie, Zhi Yang, Yangyu Tao, Bin Cui
Abstract
Embedding models have been an effective learning paradigm for high-dimensional data. However, one open issue of embedding models is that their representations (latent factors) often result in large parameter space. We observe that existing distributed training frameworks face a scalability issue of embedding models since updating and retrieving the shared embedding parameters from servers usually dominates the training cycle. In this paper, we propose HET, a new system framework that significantly improves the scalability of huge embedding model training. We embrace skewed popularity distributions of embeddings as a performance opportunity and leverage it to address the communication bottleneck with an embedding cache. To ensure consistency across the caches, we incorporate a new consistency model into HET design, which provides fine-grained consistency guarantees on a per-embedding basis. Compared to previous work that only allows staleness for read operations, HET also utilizes staleness for write operations. Evaluations on six representative tasks show that HET achieves up to 88% embedding communication reductions and up to 20.68×performance speedup over the state-of-the-art baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers29
- Galvatron: Efficient Transformer Training over Multiple GPUs Using Automatic ParallelismXupeng Miao, Yujie Wang, Youhe Jiang, Chunan Shi et al.VLDB 2023 · 113 citations
- BlindFL: Vertical Federated Machine Learning without Peeking into Your DataFangcheng Fu, Huanran Xue, Yong Cheng, Yangyu Tao et al.SIGMOD 2022 · 53 citations
- SDPipe: A Semi-Decentralized Framework for Heterogeneity-aware Pipeline-parallel TrainingXupeng Miao, Yining Shi, Zhi Yang, Bin Cui et al.VLDB 2023 · 48 citations
- FlexMoE: Scaling Large-scale Sparse Pre-trained Model Training via Dynamic Device PlacementXiaonan Nie, Xupeng Miao, Zilong Wang, Zichao Yang et al.SIGMOD 2023 · 40 citations
- Towards Communication-efficient Vertical Federated Learning Training via Cache-enabled Local UpdateFangcheng Fu, Xupeng Miao, Jiawei Jiang, Huanran Xue et al.VLDB 2022 · 31 citations
Builds on3
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
- Heterogeneity-Aware Distributed Machine Learning Training via Partial ReduceXupeng Miao, Xiaonan Nie, Yingxia Shao, Zhi Yang et al.SIGMOD 2021 · 64 citations
- Kraken: memory-efficient continual learning for large-scale real-time recommendationsMinhui Xie, Kai Ren, Youyou Lu, Guangxu Yang et al.SC 2020 · 38 citations
Related papers
- HET-GMP: A Graph-based System Approach to Scaling Large Embedding Model TrainingXupeng Miao, Yining Shi, Hailin Zhang, Xin Zhang et al.SIGMOD 2022 · 24 citations
- Accelerating Neural Recommendation Training with Embedding SchedulingChaoliang Zeng, Xudong Liao, Xiaodian Cheng, Han Tian et al.NSDI 2024 · 16 citations
- HET-KG: Communication-Efficient Knowledge Graph Embedding Training via Hotness-Aware CacheSicong Dong, Xupeng Miao, Pengkai Liu, Xin Wang et al.ICDE 2022 · 13 citations
- FEC: Efficient Deep Recommendation Model Training with Flexible Embedding CommunicationKaihao Ma, Xiao Yan, Zhenkun Cai, Yuzhen Huang et al.SIGMOD 2023 · 8 citations
- DistVec: Efficient Distributed Machine Learning in Parallel Database SystemsXinyi Zhang, Liangzu Liu, Xupeng Miao, Yinjun Wu et al.ICDE 2026
