Flare: flexible in-network allreduce
Daniele De Sensi, Salvatore Di Girolamo, Saleh Ashkboos, Shigang Li, Torsten Hoefler
摘要
The allreduce operation is one of the most commonly used communication routines in distributed applications. To improve its bandwidth and to reduce network traffic, this operation can be accelerated by offloading it to network switches, that aggregate the data received from the hosts, and send them back the aggregated result. However, existing solutions provide limited customization opportunities and might provide suboptimal performance when dealing with custom operators and data types, with sparse data, or when reproducibility of the aggregation is a concern. To deal with these problems, in this work we design a flexible programmable switch by using as a building block PsPIN, a RISC-V architecture implementing the sPIN programming model. We then design, model, and analyze different algorithms for executing the aggregation on this architecture, showing performance improvements compared to state-of-the-art approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Efficient sparse collective communication and its application to accelerate distributed deep learningJiawei Fei, Chen-Yu Ho, Atal Narayan Sahu, Marco Canini 等SIGCOMM 2021 · 被引用 120 次
- Using trio: juniper networks' programmable chipset - for emerging in-network applicationsMingran Yang, Alex Baban, Valery Kugel, Jeff Libby 等SIGCOMM 2022 · 被引用 57 次
- Swing: Short-cutting Rings for Higher Bandwidth AllreduceDaniele De Sensi, Tommaso Bonato, David Saam, Torsten HoeflerNSDI 2024 · 被引用 48 次
- In-Network Aggregation with Transport Transparency for Distributed TrainingShuo Liu, Qiaoling Wang, Junyi Zhang, Wenfei Wu 等ASPLOS 2023 · 被引用 46 次
- Themis: a network bandwidth-aware collective scheduling policy for distributed training of DL modelsSaeed Rashidi, William Won, Sudarshan Srinivasan, Srinivas Sridharan 等ISCA 2022 · 被引用 40 次
它引用的顶会 Paper8
- ATP: In-network Aggregation for Multi-tenant LearningChonLam Lao, Yanfang Le, Kshiteej Mahajan, Yixi Chen 等NSDI 2021 · 被引用 359 次
- An in-depth analysis of the slingshot interconnectDaniele De Sensi, Salvatore Di Girolamo, Kim H. McMahon, Duncan Roweth 等SC 2020 · 被引用 122 次
- Efficient sparse collective communication and its application to accelerate distributed deep learningJiawei Fei, Chen-Yu Ho, Atal Narayan Sahu, Marco Canini 等SIGCOMM 2021 · 被引用 120 次
- An In-Network Architecture for Accelerating Shared-Memory Multiprocessor CollectivesBenjamin Klenk, Nan Jiang, Greg Thorson, Larry DennisonISCA 2020 · 被引用 67 次
- Taming unbalanced training workloads in deep learning with partial collective operationsShigang Li, Tal Ben-Nun, Salvatore Di Girolamo, Dan Alistarh 等PPoPP 2020 · 被引用 52 次
相关 Paper
- A RISC-V in-network accelerator for flexible high-performance low-power packet processingSalvatore Di Girolamo, Andreas Kurth, Alexandru Calotoiu, Thomas Benz 等ISCA 2021 · 被引用 37 次
- Near-Optimal Wafer-Scale ReducePiotr Luczynski, Lukas Gianinazzi, Patrick Iff, Leighton Wilson 等HPDC 2024 · 被引用 7 次
- A Generic Service to Provide In-Network Aggregation for Key-Value StreamsYongchao He, Wenfei Wu, Yanfang Le, Ming Liu 等ASPLOS 2023 · 被引用 37 次
- Scaling Distributed Machine Learning with In-Network AggregationAmedeo Sapio, Marco Canini, Chen-Yu Ho, Jacob Nelson 等NSDI 2021
- RED: Distributed Program Deployment for Resource-aware Programmable SwitchesXingxin Jia, Fuliang Li, Songlin Chen, Chengxi Gao 等INFOCOM 2023 · 被引用 6 次
