Efficient and Flexible Datapaths for Fine-Grained Rack-Scale Interconnects with Elastic QP
Chenxingyu Zhao, Yibo Wu, Hongtao Zhang, Jaehong Min, Ming Liu, Arvind Krishnamurthy
摘要
Rack-scale interconnects serve as critical datapaths for emerging communication-intensive systems to scale up. Innovative solutions for this datapath are rising at a rapid pace, especially those based on Ethernet. However, existing hardware-based solutions, such as RDMA, face performance issues, particularly for small-message memory access, and suffer from the inflexibility of hardware-fixed processing. The community is actively pursuing efficient, flexible, and cost-effective rack-scale datapaths.
In this work, we propose Software-Interposed Datapath (SID), an efficient, software-flexible, and low-cost solution for rack-scale interconnects, particularly optimized for fine-grained memory access. Improving small-message efficiency is a well-known challenge, and software involvement for flexibility seems to amplify the performance hurdle further. SID boosts performance by exploiting one insight: existing NICs primarily rely on Queue Pair (QP)-level parallelism, but underutilize intra-QP Work Queue Element (WQE)level parallelism. Harnessing parallelism is non-trivial, especially at the WQE-level, due to ordering semantics and request dispatching. The key technique is our Elastic QP data structure built atop the on-NIC datapath processors, which realizes ordered intra-QP parallelism while minimizing coordination overhead. Regarding flexibility, SID supports extensible operation sets that comply with the OpenSHMEM model for ML/HPC workloads and the Message Queue model for cloud service workloads. Regarding cost efficiency, SID is built on top of commodity components such as Ethernet, PCIe, and datapath cores of NVIDIA ConnectX-8 and BlueField-3 NICs. Evaluation shows that SID achieves up to 11.03x higher rates for small messages than RDMA-based baselines and supports both CPU and GPU-Direct operations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper74
- SIGMA: A Sparse and Irregular GEMM Accelerator with Flexible Interconnects for DNN TrainingEric Qin, Ananda Samajdar, Hyoukjun Kwon, Vineet Nadella 等HPCA 2020 · 被引用 490 次
- Alibaba HPN: A Data Center Network for Large Language Model TrainingKun Qian, Yongqing Xi, Jiamin Cao, Jiaqi Gao 等SIGCOMM 2024 · 被引用 173 次
- Can far memory improve job throughput?Emmanuel Amaro, Christopher Branner-Augmon, Zhihong Luo, Amy Ousterhout 等EuroSys 2020 · 被引用 163 次
- SRNIC: A Scalable Architecture for RDMA NICsZilong Wang, Layong Luo, Qingsong Ning, Chaoliang Zeng 等NSDI 2023 · 被引用 154 次
- Empowering Azure Storage with RDMAWei Bai, Shanim Sainul Abdeen, Ankit Agrawal, Krishan Kumar Attre 等NSDI 2023 · 被引用 117 次
相关 Paper
- DRack: A CXL-Disaggregated Rack Architecture to Boost Inter-Rack CommunicationXu Zhang, Ke Liu, Yuan Hui, Xiaolong Zheng 等USENIX ATC 2025 · 被引用 5 次
- Efficient Remote Memory Ordering for Non-Coherent SystemsWei Siew Liew, Md Ashfaqur Rahaman, Adarsh Patil, Ryan Stutsman 等ASPLOS 2026
- White-Boxing RDMA with Packet-Granular Software ControlChenxingyu Zhao, Jaehong Min, Ming Liu, Arvind KrishnamurthyNSDI 2025 · 被引用 28 次
- Maximizing the Benefit of RDMA at End HostsXiaoliang Wang, Hexiang Song, Cam-Tu Nguyen, Dongxu Cheng 等INFOCOM 2021 · 被引用 8 次
- Symphony: Enhancing RDMA Connection Scalability through Sender-Receiver CoordinationYuxuan Hu, Jiao Zhang, Dexuan Liao, Xianyu HuangINFOCOM 2026
