Birds of a Feather Flock Together: Scaling RDMA RPCs with Flock
Sumit Kumar Monga, Sanidhya Kashyap, Changwoo Min
Abstract
RDMA-capable networks are gaining traction with datacenter deployments due to their high throughput, low latency, CPU efficiency, and advanced features, such as remote memory operations. However, efficiently utilizing RDMA capability in a common setting of high fan-in, fan-out asymmetric network topology is challenging. For instance, using RDMA programming features comes at the cost of connection scalability, which does not scale with increasing cluster size. To address that, several works forgo some RDMA features by only focusing on conventional RPC APIs.
In this work, we strive to exploit the full capability of RDMA, while scaling the number of connections regardless of the cluster size. We present Flock, a communication framework for RDMA networks that uses hardware provided reliable connection. Using a partially shared model, Flock departs from the conventional RDMA design by enabling connection sharing among threads, which provides significant performance improvements contrary to the widely held belief that connection sharing deteriorates performance. At its core, Flock uses a connection handle abstraction for connection multiplexing; a new coalescing-based synchronization approach for efficient network utilization; and a loadcontrol mechanism for connections with symbiotic send-recv scheduling, which reduces the synchronization overheads associated with connection sharing along with ensuring fair utilization of network connections. We demonstrate the benefits for a distributed transaction processing system and an in-memory index, where it outperforms other RPC systems by up to 88% and 50%, respectively, with significant reductions in median and tail latency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 740c0695-5dd4-4078-ac48-99a5f293da80Cited by top-tier papers16
- Understanding RDMA Microarchitecture Resources for Performance IsolationXinhao Kong, Jingrong Chen, Wei Bai, Yechen Xu et al.NSDI 2023 · 81 citations
- Characterizing Off-path SmartNIC for Accelerating Distributed SystemsXingda Wei, Rongxin Cheng, Yuhan Yang, Rong Chen et al.OSDI 2023 · 68 citations
- DINT: Fast In-Kernel Distributed Transactions with eBPFYang Zhou, Xingyu Xiang, Matthew Kiley, Sowmya Dharanipragada et al.NSDI 2024 · 40 citations
- DEEPSERVE: Serverless Large Language Model Serving at ScaleJunhao Hu, Jiang Xu, Zhixia Liu, Yulong He et al.USENIX ATC 2025 · 38 citations
- Remote Procedure Call as a Managed System ServiceJingrong Chen, Yongji Wu, Shihan Lin, Yechen Xu et al.NSDI 2023 · 30 citations
Builds on3
- 1RMA: Re-envisioning Remote Memory Access for Multi-tenant DatacentersArjun Singhvi, Aditya Akella, Dan Gibson, Thomas F. Wenisch et al.SIGCOMM 2020 · 70 citations
- Overload Control for µs-scale RPCs with BreakwaterInho Cho, Ahmed Saeed, Joshua Fried, Seo Jin Park et al.OSDI 2020 · 61 citations
- HydraList: A Scalable In-Memory Index Using Asynchronous Updates and Partial ReplicationAjit Mathew, Changwoo MinVLDB 2020 · 22 citations
Related papers
- Practical and Scalable RDMA Connection Sharing for HPC WorkloadYuejie Wang, Tuo Fang, Biyu Peng, Yang Cheng et al.EuroSys 2026
- FARLock: Asymmetric RDMA Locking Made FairYuehao Hu, Jiatang Zhou, Tianzheng Wang, Keval VoraOSDI 2026
- Fast Distributed Transactions for RDMA-based Disaggregated MemoryHaodi Lu, Haikun Liu, Yujian Zhang, Zhuohui Duan et al.USENIX ATC 2025 · 9 citations
- SF-STACK: Streamlining RDMA for Heterogeneous Telecom StorageJian Tang, Wenming Zheng, Xiaoliang Wang, Xiaoping Fan et al.INFOCOM 2026
- Scalable RDMA-accelerated Distributed Locks with Shared Stream AbstractionMiao Cai, Junru Shen, Xiaojian Liao, Rong Gu et al.EuroSys 2026
