RDMA over Ethernet for Distributed Training at Meta Scale
Adithya Gangidi, Rui Miao, Shengbao Zheng, Sai Jayesh Bondu, Guilherme Loch Waltrick Goes, Hany Morsy, Rohit Puri, Mohammad Riftadi, Ashmitha Jeevaraj Shetty, Jingyi Yang, Shuqiang Zhang, Mikel Jimenez Fernandez
2024Year
171Citations
45Top-tier citations
Abstract
The rapid growth in both computational density and scale in AI models in recent years motivates the construction of an efficient and reliable dedicated network infrastructure. This paper presents the design, implementation, and operation of Meta's Remote Direct Memory Access over Converged Ethernet (RoCE) networks for distributed AI training.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 494e6e1c-a77c-4916-ad75-2f2fbf69fef7Cited by top-tier papers45
- White-Boxing RDMA with Packet-Granular Software ControlChenxingyu Zhao, Jaehong Min, Ming Liu, Arvind KrishnamurthyNSDI 2025 · 28 citations
- Holmes: Localizing Irregularities in LLM Training with Mega-scale GPU ClustersZhiyi Yao, Pengbo Hu, Congcong Miao, Xuya Jia et al.NSDI 2025 · 23 citations
- Revisiting RDMA Reliability for Lossy FabricsWenxue Li, Xiangzhou Liu, Yunxuan Zhang, Zihao Wang et al.SIGCOMM 2025 · 21 citations
- Fire-Flyer AI-HPC: A Cost-Effective Software-Hardware Co-Design for Deep LearningWei An, Xiao Bi, Guanting Chen, Shanhuang Chen et al.SC 2024 · 20 citations
- eTran: Extensible Kernel Transport with eBPFZhongjie Chen, Qingkai Meng, ChonLam Lao, Yifan Liu et al.NSDI 2025 · 17 citations
Related papers
- Enabling AI Network Cross-Layer Design and Operations with Arcadia: A Simulation Platform at ScaleZhaodong Wang, Satyajeet Singh Ahuja, Xu Zhang, Max Noormohammadpour et al.NSDI 2026 · 3 citations
- Vela: A Virtualized LLM Training System with GPU Direct RoCEApoorve Mohan, Robert Walkup, Bengi Karacali, Ming-Hung Chen et al.ASPLOS 2025 · 3 citations
- STORM: Enabling Traffic Scheduling for RDMAJichun Wu, Ran Shu, Gianni Antichi, Yongqiang Xiong et al.SIGCOMM 2026
- Matryoshka: Realizing Hyperscale Data Center Network Design for the AI EraYan Cai, Jialong Li, Kutalmis Akpinar, Tianxiang Li et al.NSDI 2026 · 3 citations
- An In-Network Architecture for Accelerating Shared-Memory Multiprocessor CollectivesBenjamin Klenk, Nan Jiang, Greg Thorson, Larry DennisonISCA 2020 · 67 citations
