Unlocking the Power of Inline Floating-Point Operations on Programmable Switches
Yifan Yuan, Omar Alama, Jiawei Fei, Jacob Nelson, Dan R. K. Ports, Amedeo Sapio, Marco Canini, Nam Sung Kim
摘要
The advent of switches with programmable dataplanes has enabled the rapid development of new network functionality, as well as providing a platform for acceleration of a broad range of application-level functionality. However, existing switch hardware was not designed with application acceleration in mind, and thus applications requiring operations or datatypes not used in traditional network protocols must resort to expensive workarounds. Applications involving floating point data, including distributed training for machine learning and distributed query processing, are key examples.
In this paper, we propose FPISA, a floating point representation designed to work efficiently in programmable switches. We first implement FPISA on an Intel Tofino switch, but find that it has limitations that impact throughput and accuracy. We then propose hardware changes to address these limitations based on the open-source Banzai switch architecture, and synthesize them in a 15-nm standard-cell library to demonstrate their feasibility. Finally, we use FPISA to implement accelerators for training for machine learning and for query processing, and evaluate their performance on a switch implementing our changes using emulation. We find that FPISA allows distributed training to use 25-75% fewer CPU cores and provide up to 85.9% better throughput in a CPU-constrained environment than SwitchML. For distributed query processing with floating point data, FPISA enables up to 2.7× better throughput than Spark.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Using trio: juniper networks' programmable chipset - for emerging in-network applicationsMingran Yang, Alex Baban, Valery Kugel, Jeff Libby 等SIGCOMM 2022 · 被引用 57 次
- In-Network Aggregation with Transport Transparency for Distributed TrainingShuo Liu, Qiaoling Wang, Junyi Zhang, Wenfei Wu 等ASPLOS 2023 · 被引用 46 次
- Leo: Online ML-based Traffic Classification at Multi-Terabit Line RateSyed Usman Jafri, Sanjay G. Rao, Vishal Shrivastav, Mohit TawarmalaniNSDI 2024 · 被引用 46 次
- THC: Accelerating Distributed Deep Learning Using Tensor Homomorphic CompressionMinghao Li, Ran Ben Basat, Shay Vargaftik, ChonLam Lao 等NSDI 2024 · 被引用 44 次
- A Generic Service to Provide In-Network Aggregation for Key-Value StreamsYongchao He, Wenfei Wu, Yanfang Le, Ming Liu 等ASPLOS 2023 · 被引用 37 次
它引用的顶会 Paper15
- A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU ClustersYimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi 等OSDI 2020 · 被引用 390 次
- ATP: In-network Aggregation for Multi-tenant LearningChonLam Lao, Yanfang Le, Kshiteej Mahajan, Yixi Chen 等NSDI 2021 · 被引用 359 次
- PINT: Probabilistic In-band Network TelemetryRan Ben Basat, Sivaramakrishnan Ramanathan, Yuliang Li, Gianni Antichi 等SIGCOMM 2020 · 被引用 268 次
- Pushing the Limits of Narrow Precision Inferencing at Cloud Scale with Microsoft Floating PointBita Darvish Rouhani, Daniel Lo, Ritchie Zhao, Ming Liu 等NeurIPS 2020 · 被引用 153 次
- Flow Event Telemetry on Programmable Data PlaneYu Zhou, Chen Sun, Hongqiang Harry Liu, Rui Miao 等SIGCOMM 2020 · 被引用 139 次
相关 Paper
- Scaling Distributed Machine Learning with In-Network AggregationAmedeo Sapio, Marco Canini, Chen-Yu Ho, Jacob Nelson 等NSDI 2021
- Flowrest: Practical Flow-Level Inference in Programmable Switches with Random ForestsAristide Tanyi-Jong Akem, Michele Gucciardo, Marco FioreINFOCOM 2023 · 被引用 63 次
- Multitenant In-Network Acceleration with SwitchVMSajy Khashab, Alon Rashelbach, Mark SilbersteinNSDI 2024 · 被引用 21 次
- Quark: Implementing Convolutional Neural Networks Entirely on Programmable Data PlaneMai Zhang, Lin Cui, Xiaoquan Zhang, Fung Po Tso 等INFOCOM 2025 · 被引用 17 次
- FENIX: Enabling In-Network DNN Inference with FPGA-Enhanced Programmable SwitchesXiangyu Gao, Tong Li, Yinchao Zhang, Ziqiang Wang 等NSDI 2026 · 被引用 12 次
