A Generic Service to Provide In-Network Aggregation for Key-Value Streams
Yongchao He, Wenfei Wu, Yanfang Le, Ming Liu, ChonLam Lao
摘要
Key-value stream aggregation is a common operation in distributed systems, which requires intensive computation and network resources. We propose a generic in-network aggregation service for key-value streams, ASK, to accelerate the aggregation operations in diverse distributed applications. ASK is a switch-host co-designed system, where the programmable switch provides a best-effort aggregation service, and the host runs a daemon to interact with applications. ASK makes in-depth optimization tailored to traffic characteristics, hardware restrictions, and network unreliable natures: it vectorizes multiple key-value tuples’ aggregation of one packet in one switch pipeline pass, which improves the per-host’s goodput; it develops a lightweight reliability mechanism for key-value stream’s asynchronous aggregation, which guarantees computation correctness; it designs a hot-key agnostic prioritization for key-skewed workloads, which improves the switch memory utilization. We prototype ASK and use it to support Spark and BytePS. The evaluation shows that ASK could accelerate pure key-value aggregation tasks by up to 155 times and big data jobs by 3-5 times, and be backward compatible with existing INA-empowered distributed training solutions with the same speedup.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- THC: Accelerating Distributed Deep Learning Using Tensor Homomorphic CompressionMinghao Li, Ran Ben Basat, Shay Vargaftik, ChonLam Lao 等NSDI 2024 · 被引用 44 次
- ClickINC: In-network Computing as a Service in Heterogeneous Programmable Data-center NetworksWenquan Xu, Zijian Zhang, Yong Feng, Haoyu Song 等SIGCOMM 2023 · 被引用 34 次
- LogNIC: A High-Level Performance Model for SmartNICsZerui Guo, Jiaxin Lin, Yuebin Bai, Daehyeok Kim 等MICRO 2023 · 被引用 17 次
- Building an Elastic Block Storage over EBOFs Using Shadow ViewsSheng Jiang, Ming LiuNSDI 2025 · 被引用 10 次
- Understanding and Profiling NVMe-over-TCP Using ntprofYuyuan Kang, Ming LiuNSDI 2025 · 被引用 9 次
它引用的顶会 Paper10
- A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU ClustersYimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi 等OSDI 2020 · 被引用 390 次
- ATP: In-network Aggregation for Multi-tenant LearningChonLam Lao, Yanfang Le, Kshiteej Mahajan, Yixi Chen 等NSDI 2021 · 被引用 359 次
- Efficient sparse collective communication and its application to accelerate distributed deep learningJiawei Fei, Chen-Yu Ho, Atal Narayan Sahu, Marco Canini 等SIGCOMM 2021 · 被引用 120 次
- An In-Network Architecture for Accelerating Shared-Memory Multiprocessor CollectivesBenjamin Klenk, Nan Jiang, Greg Thorson, Larry DennisonISCA 2020 · 被引用 67 次
- Harmonia: Near-Linear Scalability for Replicated Storage with In-Network Conflict DetectionHang Zhu, Zhihao Bai, Jialin Li, Ellis Michael 等VLDB 2020 · 被引用 58 次
相关 Paper
- Training Job Placement in Clusters with Statistical In-Network AggregationBohan Zhao, Wei Xu, Shuo Liu, Yang Tian 等ASPLOS 2024 · 被引用 17 次
- Spatiotemporal Sketch Disaggregation: Streaming Analytics with Heterogeneous ResourcesJonatan Langlet, Peiqing Chen, Michael Mitzenmacher, Zaoxing Liu 等ICDE 2026
- SwitchTop-k: Scaling Top-k Compression on Programmable SwitchesYijun Li, Jiawei Huang, Jingling Liu, Zhaoyi Li 等KDD 2025
- Flare: flexible in-network allreduceDaniele De Sensi, Salvatore Di Girolamo, Saleh Ashkboos, Shigang Li 等SC 2021 · 被引用 49 次
- SWARM: Replicating Shared Disaggregated-Memory Data in No TimeAntoine Murat, Clément Burgelin, Athanasios Xygkis, Igor Zablotchi 等SOSP 2024 · 被引用 2 次
