Birds of a Feather Flock Together: Scaling RDMA RPCs with Flock
Sumit Kumar Monga, Sanidhya Kashyap, Changwoo Min
摘要
RDMA-capable networks are gaining traction with datacenter deployments due to their high throughput, low latency, CPU efficiency, and advanced features, such as remote memory operations. However, efficiently utilizing RDMA capability in a common setting of high fan-in, fan-out asymmetric network topology is challenging. For instance, using RDMA programming features comes at the cost of connection scalability, which does not scale with increasing cluster size. To address that, several works forgo some RDMA features by only focusing on conventional RPC APIs.
In this work, we strive to exploit the full capability of RDMA, while scaling the number of connections regardless of the cluster size. We present Flock, a communication framework for RDMA networks that uses hardware provided reliable connection. Using a partially shared model, Flock departs from the conventional RDMA design by enabling connection sharing among threads, which provides significant performance improvements contrary to the widely held belief that connection sharing deteriorates performance. At its core, Flock uses a connection handle abstraction for connection multiplexing; a new coalescing-based synchronization approach for efficient network utilization; and a loadcontrol mechanism for connections with symbiotic send-recv scheduling, which reduces the synchronization overheads associated with connection sharing along with ensuring fair utilization of network connections. We demonstrate the benefits for a distributed transaction processing system and an in-memory index, where it outperforms other RPC systems by up to 88% and 50%, respectively, with significant reductions in median and tail latency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Understanding RDMA Microarchitecture Resources for Performance IsolationXinhao Kong, Jingrong Chen, Wei Bai, Yechen Xu 等NSDI 2023 · 被引用 81 次
- Characterizing Off-path SmartNIC for Accelerating Distributed SystemsXingda Wei, Rongxin Cheng, Yuhan Yang, Rong Chen 等OSDI 2023 · 被引用 68 次
- DINT: Fast In-Kernel Distributed Transactions with eBPFYang Zhou, Xingyu Xiang, Matthew Kiley, Sowmya Dharanipragada 等NSDI 2024 · 被引用 40 次
- DEEPSERVE: Serverless Large Language Model Serving at ScaleJunhao Hu, Jiang Xu, Zhixia Liu, Yulong He 等USENIX ATC 2025 · 被引用 38 次
- Remote Procedure Call as a Managed System ServiceJingrong Chen, Yongji Wu, Shihan Lin, Yechen Xu 等NSDI 2023 · 被引用 30 次
它引用的顶会 Paper3
- 1RMA: Re-envisioning Remote Memory Access for Multi-tenant DatacentersArjun Singhvi, Aditya Akella, Dan Gibson, Thomas F. Wenisch 等SIGCOMM 2020 · 被引用 70 次
- Overload Control for µs-scale RPCs with BreakwaterInho Cho, Ahmed Saeed, Joshua Fried, Seo Jin Park 等OSDI 2020 · 被引用 61 次
- HydraList: A Scalable In-Memory Index Using Asynchronous Updates and Partial ReplicationAjit Mathew, Changwoo MinVLDB 2020 · 被引用 22 次
相关 Paper
- Practical and Scalable RDMA Connection Sharing for HPC WorkloadYuejie Wang, Tuo Fang, Biyu Peng, Yang Cheng 等EuroSys 2026
- FARLock: Asymmetric RDMA Locking Made FairYuehao Hu, Jiatang Zhou, Tianzheng Wang, Keval VoraOSDI 2026
- Fast Distributed Transactions for RDMA-based Disaggregated MemoryHaodi Lu, Haikun Liu, Yujian Zhang, Zhuohui Duan 等USENIX ATC 2025 · 被引用 9 次
- SF-STACK: Streamlining RDMA for Heterogeneous Telecom StorageJian Tang, Wenming Zheng, Xiaoliang Wang, Xiaoping Fan 等INFOCOM 2026
- Scalable RDMA-accelerated Distributed Locks with Shared Stream AbstractionMiao Cai, Junru Shen, Xiaojian Liao, Rong Gu 等EuroSys 2026
