A Generic and Efficient Communication Framework for Message-Level In-Network Computing
Xinchen Wan, Luyang Li, Han Tian, Xudong Liao, Xinyang Huang, Chaoliang Zeng, Zilong Wang, Xinyu Yang, Ke Cheng, Qingsong Ning, Guyue Liu, Layong Luo, Kai Chen
Abstract
Message-Ievel in-network computing (MINC) emerges as a promising hardware acceleration method that utilizes accelerators to offload message-level computation and enhance application performance in the datacenter. However, the development of MIN C applications is challenging in the communication aspect due to poor portability and under-utilized resource. In this paper, we present Leo, a generic and efficient commu-nication framework for MINC. Leo facilitates portability across both application and hardware via introducing a communication path abstraction, which is capable of describing generic applications with predictable communication performance across diverse hardware. It further incorporates a built-in multi-path communication over CPU and accelerator to enhance communication effi-ciency. We have implemented a prototype of Leo and evaluated it with four case studies on testbeds covering FPGA-based, SoC-based smartNICs and GPU. Experiments show that Leo achieves genericity and efficiency across MINC applications, yielding 1.2-4.7 x speedup over baselines with negligible overhead.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e9824e65-04bc-4c5d-8f56-dfa413d902c0Cited by top-tier papers1
Ask how each one uses itBuilds on18
- A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU ClustersYimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi et al.OSDI 2020 · 390 citations
- ATP: In-network Aggregation for Multi-tenant LearningChonLam Lao, Yanfang Le, Kshiteej Mahajan, Yixi Chen et al.NSDI 2021 · 359 citations
- Swift: Delay is Simple and Effective for Congestion Control in the DatacenterGautam Kumar, Nandita Dukkipati, Keon Jang, Hassan M. G. Wassel et al.SIGCOMM 2020 · 333 citations
- SRNIC: A Scalable Architecture for RDMA NICsZilong Wang, Layong Luo, Qingsong Ning, Chaoliang Zeng et al.NSDI 2023 · 154 citations
- Lyra: A Cross-Platform Language and Compiler for Data Plane Programming on Heterogeneous ASICsJiaqi Gao, Ennan Zhai, Hongqiang Harry Liu, Rui Miao et al.SIGCOMM 2020 · 82 citations
Related papers
- Lynx: A SmartNIC-driven Accelerator-centric Architecture for Network ServersMaroun Tork, Lina Maudlej, Mark SilbersteinASPLOS 2020 · 64 citations
- NetRPC: Enabling In-Network Computation in Remote Procedure CallsBohan Zhao, Wenfei Wu, Wei XuNSDI 2023 · 23 citations
- Enabling In-Network Acceleration Over the CloudHao Wang, Decang Sun, Jinbin Hu, Kai ChenINFOCOM 2025 · 3 citations
- RoPeerTo: A Datacenter-Scale Architecture for Peer-To-Peer DMA between GPUs and FPGAsMarco Venere, Giuseppe Sorrentino, Benjamin Ramhorst, Maximilian Jakob Heer et al.EuroSys 2026 · 1 citation
- MARS: Exploiting Multi-Level Parallelism for DNN Workloads on Adaptive Multi-Accelerator SystemsGuan Shen, Jieru Zhao, Zeke Wang, Zhe Lin et al.DAC 2023 · 5 citations
