Accelerating Neural Recommendation Training with Embedding Scheduling
Chaoliang Zeng, Xudong Liao, Xiaodian Cheng, Han Tian, Xinchen Wan, Hao Wang, Kai Chen
摘要
Deep learning recommendation models (DLRM) are extensively adopted to support many online services. Typical DLRM training frameworks adopt the parameter server (PS) in CPU servers to maintain memory-intensive embedding tables, and leverage GPU workers with embedding cache to accelerate compute-intensive neural network computation and enable fast embedding lookups. However, such distributed systems suffer from significant communication overhead caused by the embedding transmissions between workers and PS. Prior work reduces the number of cache embedding transmissions by compromising model accuracy, including oversampling hot embeddings or applying staleness-tolerant updates.
This paper reveals that many of such transmissions can be avoided given the predictability and infrequency natures of in-cache embedding accesses in distributed training. Based on this observation, we explore a new direction to accelerate distributed DLRM training without compromising model accuracy, i.e., embedding scheduling-with the core idea of proactively determining "where embeddings should be trained" and "which embeddings should be synchronized" to increase the cache hit rate and decrease unnecessary updates, thus achieving a low communication overhead. To realize this idea, we design Herald, a real-time embedding scheduler consisting of two main components: an adaptive location-aware inputs allocator to determine where embeddings should be trained and an optimal communication plan generator to determine which embeddings should be synchronized. Our experiments with real-world workloads show that Herald reduces 48%-89% embedding transmissions, leading up to 2.11× and up to 1.61× better performance with TCP and RDMA, respectively, over 100 Gbps Ethernet for end-to-end DLRM training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- GREEN: Carbon-efficient Resource Scheduling for Machine Learning ClustersKaiqiang Xu, Decang Sun, Han Tian, Junxue Zhang 等NSDI 2025 · 被引用 23 次
- FusedRec: Fused Embedding Communication for Distributed Recommendation Training on GPUsXuanteng Huang, Fan Li, Riyang Hu, Jianchang Zhang 等AAAI 2026 · 被引用 1 次
它引用的顶会 Paper10
- A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU ClustersYimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi 等OSDI 2020 · 被引用 390 次
- Efficient sparse collective communication and its application to accelerate distributed deep learningJiawei Fei, Chen-Yu Ho, Atal Narayan Sahu, Marco Canini 等SIGCOMM 2021 · 被引用 120 次
- Accelerating Recommendation System Training by Leveraging Popular ChoicesMuhammad Adnan, Yassaman Ebrahimzadeh Maboud, Divya Mahajan, Prashant J. NairVLDB 2022 · 被引用 70 次
- HET: Scaling out Huge Embedding Model Training via Cache-enabled Distributed FrameworkXupeng Miao, Hailin Zhang, Yining Shi, Xiaonan Nie 等VLDB 2022 · 被引用 70 次
- Towards Domain-Specific Network Transport for Distributed DNN TrainingHao Wang, Han Tian, Jingrong Chen, Xinchen Wan 等NSDI 2024 · 被引用 54 次
相关 Paper
- HypeReca: Distributed Heterogeneous In-Memory Embedding Database for Training Recommender ModelsJiaao He, Shengqi Chen, Kezhao Huang, Jidong ZhaiUSENIX ATC 2025 · 被引用 2 次
- FEC: Efficient Deep Recommendation Model Training with Flexible Embedding CommunicationKaihao Ma, Xiao Yan, Zhenkun Cai, Yuzhen Huang 等SIGMOD 2023 · 被引用 8 次
- Bagpipe: Accelerating Deep Recommendation Model TrainingSaurabh Agarwal, Chengpo Yan, Ziyi Zhang, Shivaram VenkataramanSOSP 2023 · 被引用 17 次
- Heterogeneous Acceleration Pipeline for Recommendation System TrainingMuhammad Adnan, Yassaman Ebrahimzadeh Maboud, Divya Mahajan, Prashant J. NairISCA 2024 · 被引用 11 次
- Hybrid Embedding Framework for Memory-Efficient Recommendation SystemsSeung Jin Yang, Hyuk-Jae Lee, Chae-Eun RheeDAC 2025
