TRACI: Network Acceleration of Input-Dynamic Communication for Large-Scale Deep Learning Recommendation Model
Guyue Huang, Hao Li, Le Qin, Jiayi Huang, Yangwook Kang, Yufei Ding, Yuan Xie
摘要
Large-scale deep learning recommendation models (DLRMs) rely on embedding layers with terabyte-scale embedding tables, which present significant challenges to memory capacity. In addition, these embedding layers exhibit sparse and random data access patterns, which demand high memory bandwidth. Multi-GPU systems provide a promising solution, allowing for the scaling of both memory and aggregated bandwidth. However, network communication bandwidth becomes a bottleneck for multi-GPU DLRM systems. Overcoming the communication bottleneck is crucial to unlocking the potential of multi-GPU systems for efficient and high-performance DLRM training.
This paper introduces TRACI, an in-network acceleration architecture designed to optimize the communication operator in embedding layers: Aggregation. While in-network acceleration has proven successful for the All-Reduce communication collective, existing solutions do not directly apply to Aggregation due to two key challenges. Firstly, existing multi-GPU shared memory operations are designed for point-to-point communication and do not allow the network to proactively optimize communication. Secondly, in Aggregation, data transfer patterns are dynamic and dependent on input, demanding the network to dynamically discover and exploit message connections on-the-fly. To address these challenges, we propose a solution that involves a novel network transaction and switch hardware design. We introduce a new network transaction that augments messages with input reuse and output reuse identifications, and can empower the network to proactively reduce
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Disdp: Disaggregating Compute, Network, and Storage for Model-Sharded Data-Parallel TrainingMo Sun, Zihan Yang, Changyue Liao, Yingtao Li 等ISCA 2026 · 被引用 1 次
- EPIC: Abstraction and Polymorphism of In-Network Collectives on EthernetYitao Yuan, Jianglong Nie, Tianyu Bai, Ruizhe Zhou 等SIGCOMM 2026 · 被引用 1 次
- Accelerating MoE with Dynamic In-Switch Computing on Multi-GPUsQijun Zhang, Chen Zhang, Zhuoshan Zhou, Haibo Wang 等ISCA 2026
它引用的顶会 Paper12
- RecNMP: Accelerating Personalized Recommendation with Near-Memory ProcessingLiu Ke, Udit Gupta, Benjamin Youngjae Cho, David Brooks 等ISCA 2020 · 被引用 235 次
- Pegasus: Tolerating Skewed Workloads in Distributed Storage with In-Network Coherence DirectoriesJialin Li, Jacob Nelson, Ellis Michael, Xin Jin 等OSDI 2020 · 被引用 96 次
- Concordia: Distributed Shared Memory with In-Network Cache CoherenceQing Wang, Youyou Lu, Erci Xu, Junru Li 等FAST 2021 · 被引用 74 次
- Accelerating Recommendation System Training by Leveraging Popular ChoicesMuhammad Adnan, Yassaman Ebrahimzadeh Maboud, Divya Mahajan, Prashant J. NairVLDB 2022 · 被引用 70 次
- An In-Network Architecture for Accelerating Shared-Memory Multiprocessor CollectivesBenjamin Klenk, Nan Jiang, Greg Thorson, Larry DennisonISCA 2020 · 被引用 67 次
相关 Paper
- HypeReca: Distributed Heterogeneous In-Memory Embedding Database for Training Recommender ModelsJiaao He, Shengqi Chen, Kezhao Huang, Jidong ZhaiUSENIX ATC 2025 · 被引用 2 次
- UpDLRM: Accelerating Personalized Recommendation using Real-World PIM ArchitectureSitian Chen, Haobin Tan, Amelie Chi Zhou, Yusen Li 等DAC 2024 · 被引用 9 次
- Accelerating Neural Recommendation Training with Embedding SchedulingChaoliang Zeng, Xudong Liao, Xiaodian Cheng, Han Tian 等NSDI 2024 · 被引用 16 次
- Accelerating Distributed DLRM Training with Optimized TT Decomposition and Micro-BatchingWeihu Wang, Yaqi Xia, Donglin Yang, Xiaobo Zhou 等SC 2024 · 被引用 3 次
- EL-Rec: Efficient Large-Scale Recommendation Model Training via Tensor-Train Embedding TableZheng Wang, Yuke Wang, Boyuan Feng, Dheevatsa Mudigere 等SC 2022 · 被引用 15 次
