1Pipe: scalable total order communication in data center networks
Bojie Li, Gefei Zuo, Wei Bai, Lintao Zhang
Abstract
This paper proposes 1Pipe, a novel communication abstraction that enables different receivers to process messages from senders in a consistent total order. More precisely, 1Pipe provides both unicast and scattering (i.e., a group of messages to different destinations) in a causally and totally ordered manner. 1Pipe provides a best effort service that delivers each message at most once, as well as a reliable service that guarantees delivery and provides restricted atomic delivery for each scattering. 1Pipe can simplify and accelerate many distributed applications, e.g., transactional key-value stores, log replication, and distributed data structures.
We propose a scalable and efficient method to implement 1Pipe inside data centers. To achieve total order delivery in a scalable manner, 1Pipe separates the bookkeeping of order information from message forwarding, and distributes the work to each switch and host. 1Pipe aggregates order information using in-network computation at switches. This forms the "control plane" of the system. On the "data plane", 1Pipe forwards messages in the network as usual and reorders them at the receiver based on the order information.
Evaluation on a 32-server testbed shows that 1Pipe achieves scalable throughput (80M messages per second per host) and low latency (10𝜇s) with little CPU and network overhead. 1Pipe achieves linearly scalable throughput and low latency in transactional keyvalue store, TPC-C, remote data structures, and replication that outperforms traditional designs by 2∼20x.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bf9498cd-7987-4ae1-91b7-271635c1371eCited by top-tier papers3
- Hydra: Serialization-Free Network Ordering for Strongly Consistent Distributed ApplicationsInho Choi, Ellis Michael, Yunfan Li, Dan R. K. Ports et al.NSDI 2023 · 21 citations
- Graham: Synchronizing Clocks by Leveraging Local Clock PropertiesAli Najafi, Michael WeiNSDI 2022
- Switch: Asynchronous Metadata Updating for Distributed Storage with in-Network Data VisibilityJunru Li, Qing Wang, Zhe Yang, Shuo Liu et al.ICDE 2026
Builds on6
- Swift: Delay is Simple and Effective for Congestion Control in the DatacenterGautam Kumar, Nandita Dukkipati, Keon Jang, Hassan M. G. Wassel et al.SIGCOMM 2020 · 333 citations
- Aeolus: A Building Block for Proactive Transport in DatacentersShuihai Hu, Wei Bai, Gaoxiong Zeng, Zilong Wang et al.SIGCOMM 2020 · 138 citations
- Sundial: Fault-tolerant Clock Synchronization for DatacentersYuliang Li, Gautam Kumar, Hema Hariharan, Hassan M. G. Wassel et al.OSDI 2020 · 66 citations
- NetLock: Fast, Centralized Lock Management Using Programmable SwitchesZhuolong Yu, Yiwen Zhang, Vladimir Braverman, Mosharaf Chowdhury et al.SIGCOMM 2020 · 60 citations
- One More Config is Enough: Saving (DC)TCP for High-speed Extremely Shallow-buffered DatacentersWei Bai, Shuihai Hu, Kai Chen, Kun Tan et al.INFOCOM 2020 · 34 citations
Related papers
- Fast ACS: Low-Latency File-Based Ordered Message Delivery at ScaleSushant Kumar Gupta, Anil Raghunath Iyer, Chang Yu, Neel Bagora et al.USENIX ATC 2025
- Scalog: Seamless Reconfiguration and Total Order in a Scalable Shared LogCong Ding, David Chu, Evan Zhao, Xiang Li et al.NSDI 2020 · 52 citations
- TawPipe: Topology-Aware Weight Pipeline Parallelism for Accelerating Long-Context Large Models TrainingHouming Wu, Ling ChenAAAI 2026
- Birds of a Feather Flock Together: Scaling RDMA RPCs with FlockSumit Kumar Monga, Sanidhya Kashyap, Changwoo MinSOSP 2021 · 34 citations
- Chop Chop: Byzantine Atomic Broadcast to the Network LimitMartina Camaioni, Rachid Guerraoui, Matteo Monti, Pierre-Louis Roman et al.OSDI 2024 · 9 citations
