SC2021Top-tier venue
Flare: flexible in-network allreduce
Daniele De Sensi, Salvatore Di Girolamo, Saleh Ashkboos, Shigang Li, Torsten Hoefler
Abstract
The allreduce operation is one of the most commonly used communication routines in distributed applications. To improve its bandwidth and to reduce network traffic, this operation can be accelerated by offloading it to network switches, that aggregate the data received from the hosts, and send them back the aggregated result. However, existing solutions provide limited customization opportunities and might provide suboptimal performance when dealing with custom operators and data types, with sparse data, or when reproducibility of the aggregation is a concern. To deal with these problems, in this work we design a flexible programmable switch by using as a building block PsPIN, a RISC-V architecture implementing the sPIN programming model. We then design, model, and analyze different algorithms for executing the aggregation on this architecture, showing performance improvements compared to state-of-the-art approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 813e9e39-22fc-4f1f-8184-57f997dce126Cited by top-tier papers18
- Efficient sparse collective communication and its application to accelerate distributed deep learningJiawei Fei, Chen-Yu Ho, Atal Narayan Sahu, Marco Canini et al.SIGCOMM 2021 · 120 citations
- Using trio: juniper networks' programmable chipset - for emerging in-network applicationsMingran Yang, Alex Baban, Valery Kugel, Jeff Libby et al.SIGCOMM 2022 · 57 citations
- Swing: Short-cutting Rings for Higher Bandwidth AllreduceDaniele De Sensi, Tommaso Bonato, David Saam, Torsten HoeflerNSDI 2024 · 48 citations
- In-Network Aggregation with Transport Transparency for Distributed TrainingShuo Liu, Qiaoling Wang, Junyi Zhang, Wenfei Wu et al.ASPLOS 2023 · 46 citations
- Themis: a network bandwidth-aware collective scheduling policy for distributed training of DL modelsSaeed Rashidi, William Won, Sudarshan Srinivasan, Srinivas Sridharan et al.ISCA 2022 · 40 citations
Builds on8
- ATP: In-network Aggregation for Multi-tenant LearningChonLam Lao, Yanfang Le, Kshiteej Mahajan, Yixi Chen et al.NSDI 2021 · 359 citations
- An in-depth analysis of the slingshot interconnectDaniele De Sensi, Salvatore Di Girolamo, Kim H. McMahon, Duncan Roweth et al.SC 2020 · 122 citations
- Efficient sparse collective communication and its application to accelerate distributed deep learningJiawei Fei, Chen-Yu Ho, Atal Narayan Sahu, Marco Canini et al.SIGCOMM 2021 · 120 citations
- An In-Network Architecture for Accelerating Shared-Memory Multiprocessor CollectivesBenjamin Klenk, Nan Jiang, Greg Thorson, Larry DennisonISCA 2020 · 67 citations
- Taming unbalanced training workloads in deep learning with partial collective operationsShigang Li, Tal Ben-Nun, Salvatore Di Girolamo, Dan Alistarh et al.PPoPP 2020 · 52 citations
Related papers
- A RISC-V in-network accelerator for flexible high-performance low-power packet processingSalvatore Di Girolamo, Andreas Kurth, Alexandru Calotoiu, Thomas Benz et al.ISCA 2021 · 37 citations
- Near-Optimal Wafer-Scale ReducePiotr Luczynski, Lukas Gianinazzi, Patrick Iff, Leighton Wilson et al.HPDC 2024 · 7 citations
- A Generic Service to Provide In-Network Aggregation for Key-Value StreamsYongchao He, Wenfei Wu, Yanfang Le, Ming Liu et al.ASPLOS 2023 · 37 citations
- Scaling Distributed Machine Learning with In-Network AggregationAmedeo Sapio, Marco Canini, Chen-Yu Ho, Jacob Nelson et al.NSDI 2021
- RED: Distributed Program Deployment for Resource-aware Programmable SwitchesXingxin Jia, Fuliang Li, Songlin Chen, Chengxi Gao et al.INFOCOM 2023 · 6 citations
