Understanding Routable PCIe Performance for Composable Infrastructures
Wentao Hou, Jie Zhang, Zeke Wang, Ming Liu
Abstract
Routable PCIe has become the predominant cluster interconnect to build emerging composable infrastructures. Empowered by PCIe non-transparent bridge devices, PCIe transactions can traverse multiple switching domains, enabling a server to elastically integrate a number of remote PCIe devices as local ones. However, it is unclear how to move data or perform communication efficiently over the routable PCIe fabric without understanding its capabilities and limitations.
This paper presents the design and implementation of rP-CIeBench 1 , a software-hardware co-designed benchmarking framework to systematically characterize the routable PCIe fabric. rPCIeBench provides flexible data communication primitives, exposes end-to-end PCIe transaction observability, and enables reconfigurable experiment deployment. Using rPCIeBench, we first analyze the communication characteristics of a routable PCIe path, quantify its performance tax, and compare it with the local PCIe link. We then use it to dissect in-fabric traffic orchestration behaviors and draw three interesting findings: approximate max-min bandwidth partition, fast end-to-end bandwidth synchronization, and interferencefree among orthogonal data paths. Finally, we encode gathered characterization insights as traffic orchestration rules and develop an edge constraints relaxing algorithm to estimate PCIe flow transmission performance over a shared fabric. We validate its accuracy and demonstrate its potential to provide an optimization guide to design efficient flow schedulers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2593afbd-0d4b-4528-a64e-3a82b74e415bCited by top-tier papers15
- CEIO: A Cache-Efficient Network I/O Architecture for NIC-CPU Data PathsBowen Liu, Xinyang Huang, Qijing Li, Zhuobin Huang et al.SIGCOMM 2025 · 10 citations
- Building an Elastic Block Storage over EBOFs Using Shadow ViewsSheng Jiang, Ming LiuNSDI 2025 · 10 citations
- Understanding and Profiling NVMe-over-TCP Using ntprofYuyuan Kang, Ming LiuNSDI 2025 · 9 citations
- Building Massive MIMO Baseband Processing on a Single-Node SupercomputerXincheng Xie, Wentao Hou, Zerui Guo, Ming LiuNSDI 2025 · 8 citations
- Understanding and Profiling CXL.mem Using PathFinderXiao Li, Zerui Guo, Yuebin Bai, Mahesh Ketkar et al.SIGCOMM 2025 · 6 citations
Builds on11
- Pond: CXL-Based Memory Pooling Systems for Cloud PlatformsHuaicheng Li, Daniel S. Berger, Lisa Hsu, Daniel Ernst et al.ASPLOS 2023 · 328 citations
- TPP: Transparent Page Placement for CXL-Enabled Tiered-MemoryHasan Al Maruf, Hao Wang, Abhishek Dhanotia, Johannes Weiner et al.ASPLOS 2023 · 255 citations
- Demystifying CXL Memory with Genuine CXL-Ready Systems and DevicesYan Sun, Yifan Yuan, Zeduo Yu, Reese Kuper et al.MICRO 2023 · 133 citations
- Overcoming the Memory Wall with CXL-Enabled SSDsShao-Peng Yang, Minjae Kim, Sanghyun Nam, Juhyung Park et al.USENIX ATC 2023 · 75 citations
- CXL-ANNS: Software-Hardware Collaborative Memory Disaggregation and Computation for Billion-Scale Approximate Nearest Neighbor SearchJunhyeok Jang, Hanjin Choi, Hanyeoreum Bae, Seungjun Lee et al.USENIX ATC 2023 · 75 citations
Related papers
- weBurst can be Harmless: Achieving Line-rate Software Traffic Shaping by Inter-flow BatchingDanfeng Shan, Shihao Hu, Yuqi Liu, Wanchun Jiang et al.INFOCOM 2023 · 2 citations
- RpcNIC: Enabling Efficient Datacenter RPC Offloading on PCIe-attached SmartNICsJie Zhang, Hongjing Huang, Xuzheng Chen, Xiang Li et al.HPCA 2025 · 6 citations
- NetTLP: A Development Platform for PCIe devices in Software Interacting with HardwareYohei Kuga, Ryo Nakamura, Takeshi Matsuya, Yuji SekiyaNSDI 2020 · 9 citations
- 1Pipe: scalable total order communication in data center networksBojie Li, Gefei Zuo, Wei Bai, Lintao ZhangSIGCOMM 2021 · 5 citations
- Powerful GPUs or Fast Interconnects: Analyzing Relational Workloads on Modern GPUsMarko Kabic, Bowen Wu, Jonas Dann, Gustavo AlonsoVLDB 2025 · 7 citations
