High-throughput and Flexible Host Networking for Accelerated Computing
Athinagoras Skiadopoulos, Zhiqiang Xie, Mark Zhao, Qizhe Cai, Saksham Agarwal, Jacob Adelmann, David Ahern, Carlo Contavalli, Michael D. Goldflam, Vitaly Mayatskikh, Raghu Raja, Daniel Walton
Abstract
Modern network hardware is able to meet the stringent bandwidth demands of applications like GPU-accelerated AI. However, existing host network stacks offer a hard tradeoff between performance (in terms of sustained throughput when compared to network hardware capacity) and flexibility (in terms of the ability to select, customize, and extend different network protocols).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 38b8f6c2-4fbe-475e-a022-7bf83cc14013Cited by top-tier papers5
- Barre: Empowering Simplified and Versatile Programmable Congestion Control in High-Speed AI ClustersYajuan Peng, Haoran Wei, Xiaolong Zhong, Junkai Huang et al.USENIX ATC 2025 · 5 citations
- Remote TCP Connection Offload and ApplicationsShuo Li, Steven W. D. Chien, Tianyi Gao, Michio HondaNSDI 2026 · 3 citations
- Presto: A Match-Action TCP Stack for the Terabit EraRajath Shashidhara, Antoine Kaufmann, Simon PeterSIGCOMM 2026 · 1 citation
- UCCL-Tran: An Extensible Software Transport Layer for GPU NetworkingYang Zhou, Zhongjie Chen, Ziming Mao, ChonLam Lao et al.OSDI 2026
- SmartNS: Enabling Line-rate and Flexible Network Stack with SmartNICXuzheng Chen, Jie Zhang, Baolin Zhu, Xueying Zhu et al.EuroSys 2026
Builds on22
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- Orca: A Distributed Serving System for Transformer-Based Generative ModelsGyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim et al.OSDI 2022 · 690 citations
- When Cloud Storage Meets RDMAYixiao Gao, Qiang Li, Lingbo Tang, Yongqing Xi et al.NSDI 2021 · 228 citations
- TopoOpt: Co-optimizing Network Topology and Parallelization Strategy for Distributed Training JobsWeiyang Wang, Moein Khazraee, Zhizhen Zhong, Manya Ghobadi et al.NSDI 2023 · 215 citations
- Caladan: Mitigating Interference at Microsecond TimescalesJoshua Fried, Zhenyuan Ruan, Amy Ousterhout, Adam BelayOSDI 2020 · 213 citations
Related papers
- <u>G</u>PU <u>i</u>nitiated <u>O</u>penSHMEM: correct and efficient intra-kernel networking for dGPUsKhaled Hamidouche, Michael LeBeanePPoPP 2020 · 17 citations
- Understanding host network stack overheadsQizhe Cai, Shubham Chaudhary, Midhul Vuppalapati, Jaehyun Hwang et al.SIGCOMM 2021 · 150 citations
- GPU-Ether: GPU-native Packet I/O for GPU Applications on Commodity EthernetChangue Jung, Suhwan Kim, Ikjun Yeom, Honguk Woo et al.INFOCOM 2021 · 7 citations
- F4T: A Fast and Flexible FPGA-based Full-stack TCP Acceleration FrameworkJunehyuk Boo, Yujin Chung, Eunjin Baek, Seongmin Na et al.ISCA 2023 · 8 citations
- Beehive: A Flexible Network Stack for Direct-Attached AcceleratorsKatie Lim, Matthew Giordano, Theano Stavrinos, Irene Zhang et al.MICRO 2024 · 4 citations
