FusedRec: Fused Embedding Communication for Distributed Recommendation Training on GPUs
Xuanteng Huang, Fan Li, Riyang Hu, Jianchang Zhang, Yuan Peng, Yang Zhou, Fangying Chen, Xianwei Zhang
Abstract
Recent years have witnessed the wide adoption of deep learning recommendation models (DLRMs) for many online services. Unlike traditional DNN training, DLRMs leverage massive embeddings to represent sparse features, which are stored in distributed GPUs following the model parallel paradigm. Existing approaches adopt deduplication to eliminate replicated embeddings involved in AltoAll transfers to avoid unnecessary communication. In our practices, we have observed that such a deduplication design exacerbates interconnect inefficiency due to the fragmented embedding transfers with reduced message sizes, hindering the performance of distributed DLRM training.
This paper introduces FusedRec, a fused embedding communication and lookup mechanism to tackle the inefficiency due to deduplication. By seeking the opportunities to fuse embeddings from multiple categories into a group, FusedRec conducts the communication in a combined shot to alleviate bandwidth under-utilization. Meanwhile, a categorical-aware hashing algorithm is integrated into FusedRec to retain the category information during lookup without extra communication. Combining with efficient unique and recovery operations, comprehensive results show FusedRec achieves a 37.8% throughput speedup in average compared to the SOTA industry implementation, without hurting the recommendation qualities of our in-house models used in online production environments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9afab984-ec11-44d7-afa3-addd2d459378Cited by top-tier papers1
Ask how each one uses itBuilds on16
- RecShard: statistical feature-based memory optimization for industry-scale neural recommendationGeet Sethi, Bilge Acun, Niket Agarwal, Christos Kozyrakis et al.ASPLOS 2022 · 65 citations
- Ekko: A Large-Scale Deep Learning Recommender System with Low-Latency Model UpdateChijun Sima, Yao Fu, Man-Kit Sit, Liyi Guo et al.OSDI 2022 · 31 citations
- Accelerating Collective Communication in Data Parallel Training across Deep Learning FrameworksJoshua Romero, Junqi Yin, Nouamane Laanait, Bing Xie et al.NSDI 2022 · 30 citations
- Training personalized recommendation systems from (GPU) scratch: look forward not backwardsYoungeun Kwon, Minsoo RhuISCA 2022 · 24 citations
- AdaEmbed: Adaptive Embedding for Large-Scale Recommendation ModelsFan Lai, Wei Zhang, Rui Liu, William Tsai et al.OSDI 2023 · 23 citations
Related papers
- Accelerating Neural Recommendation Training with Embedding SchedulingChaoliang Zeng, Xudong Liao, Xiaodian Cheng, Han Tian et al.NSDI 2024 · 16 citations
- HypeReca: Distributed Heterogeneous In-Memory Embedding Database for Training Recommender ModelsJiaao He, Shengqi Chen, Kezhao Huang, Jidong ZhaiUSENIX ATC 2025 · 2 citations
- Hybrid Embedding Framework for Memory-Efficient Recommendation SystemsSeung Jin Yang, Hyuk-Jae Lee, Chae-Eun RheeDAC 2025
- Accelerating Distributed DLRM Training with Optimized TT Decomposition and Micro-BatchingWeihu Wang, Yaqi Xia, Donglin Yang, Xiaobo Zhou et al.SC 2024 · 3 citations
- FEC: Efficient Deep Recommendation Model Training with Flexible Embedding CommunicationKaihao Ma, Xiao Yan, Zhenkun Cai, Yuzhen Huang et al.SIGMOD 2023 · 8 citations
