White-Boxing RDMA with Packet-Granular Software Control
Chenxingyu Zhao, Jaehong Min, Ming Liu, Arvind Krishnamurthy
Abstract
Driven by diverse workloads and deployments, numerous innovations emerge to customize RDMA transport, spanning congestion control, multi-tenant isolation, routing, and more. However, RDMA's hardware-offloading nature poses significant rigidity when landing these innovations. Prior workflows to deliver customizations have either waited for lengthy hardware iterations, developed bespoke hardware, or applied coarse-grained control over the black-box RDMA NIC. Despite considerable efforts, current customization workflows still lack flexibility, raw performance, and broad availability.
In this work, we advocate for White-Boxing RDMA, which provides control of the hardware transport to general-purpose software while preserving raw data path performance. To facilitate the white-boxing methodology, we design and implement Software-Controlled RDMA (SCR), a framework enabling packet-granular software control over the hardware transport. To address challenges stemming from granular control over high-speed line rates, SCR employs effective control models, boosts the efficiency of subsystems within the framework, and leverages emerging hardware capabilities. We implement SCR on the latest Nvidia BlueField-3 equipped with Datapath Accelerators, delivering a spectrum of new customizations not present in legacy RDMA transport, such as Multi-Tenant Fair Scheduler, User-Defined Congestion Control, Receiver-Driven Flow Control, and Multi-Path Routing Selection. Furthermore, we demonstrate SCR's applicability for GPU-Direct and NVMe-oF RDMA with zero modifications to machine learning or storage code.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cffdae33-eb45-4666-9dd4-ad733add5b99Cited by top-tier papers12
- Understanding and Profiling CXL.mem Using PathFinderXiao Li, Zerui Guo, Yuebin Bai, Mahesh Ketkar et al.SIGCOMM 2025 · 6 citations
- Co-Designing Traffic Control with NVMe-oF for Disaggregated Storage: A Comparative Study of Switched and Switchless SAN ArchitecturesChendong Wang, Joontaek Oh, Ming LiuNSDI 2026 · 3 citations
- SG-IOV: Socket-Granular I/O Virtualization for SmartNIC-Based Container NetworksChenxingyu Zhao, Hongtao Zhang, Jaehong Min, Shengkai Lin et al.ASPLOS 2026 · 2 citations
- Building A CSFQ-Inspired Transport for Switched CXL Memory PoolingZerui Guo, Emily Shriver, Ming LiuNSDI 2026 · 2 citations
- Efficient and Flexible Datapaths for Fine-Grained Rack-Scale Interconnects with Elastic QPChenxingyu Zhao, Yibo Wu, Hongtao Zhang, Jaehong Min et al.SIGCOMM 2026 · 1 citation
Builds on36
- MegaScale: Scaling Large Language Model Training to More Than 10, 000 GPUsZiheng Jiang, Haibin Lin, Yinmin Zhong, Qi Huang et al.NSDI 2024 · 415 citations
- When Cloud Storage Meets RDMAYixiao Gao, Qiang Li, Lingbo Tang, Yongqing Xi et al.NSDI 2021 · 228 citations
- Alibaba HPN: A Data Center Network for Large Language Model TrainingKun Qian, Yongqing Xi, Jiamin Cao, Jiaqi Gao et al.SIGCOMM 2024 · 173 citations
- RDMA over Ethernet for Distributed Training at Meta ScaleAdithya Gangidi, Rui Miao, Shengbao Zheng, Sai Jayesh Bondu et al.SIGCOMM 2024 · 171 citations
- SRNIC: A Scalable Architecture for RDMA NICsZilong Wang, Layong Luo, Qingsong Ning, Chaoliang Zeng et al.NSDI 2023 · 154 citations
Related papers
- UCCL-Tran: An Extensible Software Transport Layer for GPU NetworkingYang Zhou, Zhongjie Chen, Ziming Mao, ChonLam Lao et al.OSDI 2026
- SwCC: Software-Programmable and Per-Packet Congestion Control in RDMA EngineHongjing Huang, Jie Zhang, Xuzheng Chen, Ziyu Song et al.USENIX ATC 2025 · 4 citations
- FlexDriver: a network driver for your acceleratorHaggai Eran, Maxim Fudim, Gabi Malka, Gal Shalom et al.ASPLOS 2022 · 14 citations
- Scalable RDMA-accelerated Distributed Locks with Shared Stream AbstractionMiao Cai, Junru Shen, Xiaojian Liao, Rong Gu et al.EuroSys 2026
- PeRF: Preemption-enabled RDMA FrameworkSugi Lee, Mingyu Choi, Ikjun Yeom, Younghoon KimUSENIX ATC 2024 · 3 citations
