HET: Scaling out Huge Embedding Model Training via Cache-enabled Distributed Framework
Xupeng Miao, Hailin Zhang, Yining Shi, Xiaonan Nie, Zhi Yang, Yangyu Tao, Bin Cui
摘要
Embedding models have been an effective learning paradigm for high-dimensional data. However, one open issue of embedding models is that their representations (latent factors) often result in large parameter space. We observe that existing distributed training frameworks face a scalability issue of embedding models since updating and retrieving the shared embedding parameters from servers usually dominates the training cycle. In this paper, we propose HET, a new system framework that significantly improves the scalability of huge embedding model training. We embrace skewed popularity distributions of embeddings as a performance opportunity and leverage it to address the communication bottleneck with an embedding cache. To ensure consistency across the caches, we incorporate a new consistency model into HET design, which provides fine-grained consistency guarantees on a per-embedding basis. Compared to previous work that only allows staleness for read operations, HET also utilizes staleness for write operations. Evaluations on six representative tasks show that HET achieves up to 88% embedding communication reductions and up to 20.68×performance speedup over the state-of-the-art baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper29
- Galvatron: Efficient Transformer Training over Multiple GPUs Using Automatic ParallelismXupeng Miao, Yujie Wang, Youhe Jiang, Chunan Shi 等VLDB 2023 · 被引用 113 次
- BlindFL: Vertical Federated Machine Learning without Peeking into Your DataFangcheng Fu, Huanran Xue, Yong Cheng, Yangyu Tao 等SIGMOD 2022 · 被引用 53 次
- SDPipe: A Semi-Decentralized Framework for Heterogeneity-aware Pipeline-parallel TrainingXupeng Miao, Yining Shi, Zhi Yang, Bin Cui 等VLDB 2023 · 被引用 48 次
- FlexMoE: Scaling Large-scale Sparse Pre-trained Model Training via Dynamic Device PlacementXiaonan Nie, Xupeng Miao, Zilong Wang, Zichao Yang 等SIGMOD 2023 · 被引用 40 次
- Towards Communication-efficient Vertical Federated Learning Training via Cache-enabled Local UpdateFangcheng Fu, Xupeng Miao, Jiawei Jiang, Huanran Xue 等VLDB 2022 · 被引用 31 次
它引用的顶会 Paper3
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong 等NeurIPS 2020 · 被引用 3,935 次
- Heterogeneity-Aware Distributed Machine Learning Training via Partial ReduceXupeng Miao, Xiaonan Nie, Yingxia Shao, Zhi Yang 等SIGMOD 2021 · 被引用 64 次
- Kraken: memory-efficient continual learning for large-scale real-time recommendationsMinhui Xie, Kai Ren, Youyou Lu, Guangxu Yang 等SC 2020 · 被引用 38 次
相关 Paper
- HET-GMP: A Graph-based System Approach to Scaling Large Embedding Model TrainingXupeng Miao, Yining Shi, Hailin Zhang, Xin Zhang 等SIGMOD 2022 · 被引用 24 次
- Accelerating Neural Recommendation Training with Embedding SchedulingChaoliang Zeng, Xudong Liao, Xiaodian Cheng, Han Tian 等NSDI 2024 · 被引用 16 次
- HET-KG: Communication-Efficient Knowledge Graph Embedding Training via Hotness-Aware CacheSicong Dong, Xupeng Miao, Pengkai Liu, Xin Wang 等ICDE 2022 · 被引用 13 次
- FEC: Efficient Deep Recommendation Model Training with Flexible Embedding CommunicationKaihao Ma, Xiao Yan, Zhenkun Cai, Yuzhen Huang 等SIGMOD 2023 · 被引用 8 次
- DistVec: Efficient Distributed Machine Learning in Parallel Database SystemsXinyi Zhang, Liangzu Liu, Xupeng Miao, Yinjun Wu 等ICDE 2026
