Scaling Up Memory Disaggregated Applications with SMART
Feng Ren, Mingxing Zhang, Kang Chen, Huaxia Xia, Zuoning Chen, Yongwei Wu
Abstract
Recent developments in RDMA networks are leading to the trend of memory disaggregation. However, the performance of each compute node is still limited by the network, especially when it needs to perform a large number of concurrent fine-grained remote accesses. According to our evaluations, existing IOPS-bound disaggregated applications do not scale well beyond 32 cores, and therefore do not take full advantage of today's many-core machines.
After an in-depth analysis of the internal architecture of RNIC, we found three major scale-up bottlenecks that limit the throughput of today's disaggregated applications:
(1) implicit contention of doorbell registers, (2) cache trashing caused by excessive outstanding work requests, and (3) wasted IOPS from unsuccessful CAS retries. However, the solutions to these problems involve many low-level details that are not familiar to application developers. To ease the burden on developers, we propose Smart, an RDMA programming framework that hides the above details by providing an interface similar to one-sided RDMA verbs.
We take 44 and 16 lines of code to refactor the state-of-theart disaggregated hash table (RACE) and persistent transaction processing system (FORD) with Smart, improving their throughput by up to 132.4× and 5.2×, respectively. We have also refactored Sherman (a recent disaggregated B + Tree) with Smart and an additional speculative lookup optimization (48 lines of code changed), which changes its memory access pattern from bandwidth-bound to IOPS-bound and leads to a speedup of 2.0×. Smart is publicly available at https://github.com/madsys-dev/smart.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 41b284ca-907e-4a03-8b86-130d348b3157Cited by top-tier papers8
- ShiftLock: Mitigate One-sided RDMA Lock Contention via HandoverJian Gao, Qing Wang, Jiwu ShuFAST 2025 · 10 citations
- DMTree: Towards Efficient Tree Indexing on Disaggregated Memory via Compute-side Collaborative DesignGuoli Wei, Yongkun Li, Haoze Song, Tao Li et al.FAST 2026 · 1 citation
- Shard: A Scalable and Resize-optimized Hash Index on Disaggregated MemoryHantian Zha, Teng Ma, Baotong Lu, Yuansen Wang et al.VLDB 2026
- Efficient, Scalable, and Fair Locking on Disaggregated Memory with Decentralized CoordinationHanze Zhang, Ke Cheng, Rong Chen, Xingda Wei et al.VLDB 2026
- STORM: Enabling Traffic Scheduling for RDMAJichun Wu, Ran Shu, Gianni Antichi, Yongqiang Xiong et al.SIGCOMM 2026
Builds on25
- AIFM: High-Performance, Application-Integrated Far MemoryZhenyuan Ruan, Malte Schwarzkopf, Marcos K. Aguilera, Adam BelayOSDI 2020 · 224 citations
- Can far memory improve job throughput?Emmanuel Amaro, Christopher Branner-Augmon, Zhihong Luo, Amy Ousterhout et al.EuroSys 2020 · 163 citations
- Disaggregating Persistent Memory and Controlling Them Remotely: An Exploration of Passive Disaggregated Key-Value StoresShin-Yeh Tsai, Yizhou Shan, Yiying ZhangUSENIX ATC 2020 · 159 citations
- Rethinking software runtimes for disaggregated memoryIrina Calciu, M. Talha Imran, Ivan Puddu, Sanidhya Kashyap et al.ASPLOS 2021 · 116 citations
- One-sided RDMA-Conscious Extendible Hashing for Disaggregated MemoryPengfei Zuo, Jiazhao Sun, Liu Yang, Shuangwu Zhang et al.USENIX ATC 2021 · 113 citations
Related papers
- SMART: A High-Performance Adaptive Radix Tree for Disaggregated MemoryXuchuan Luo, Pengfei Zuo, Jiacheng Shen, Jiazhen Gu et al.OSDI 2023 · 21 citations
- Cowbird: Freeing CPUs to Compute by Offloading the Disaggregation of MemoryXinyi Chen, Liangcheng Yu, Vincent Liu, Qizhen ZhangSIGCOMM 2023 · 15 citations
- Sherman: A Write-Optimized Distributed B+Tree Index on Disaggregated MemoryQing Wang, Youyou Lu, Jiwu ShuSIGMOD 2022 · 99 citations
- FORD: Fast One-sided RDMA-based Distributed Transactions for Disaggregated Persistent MemoryMing Zhang, Yu Hua, Pengfei Zuo, Lurong LiuFAST 2022 · 97 citations
- SmartDS: Middle-Tier-centric SmartNIC Enabling Application-aware Message Split for Disaggregated Block StorageJie Zhang, Hongjing Huang, Lingjun Zhu, Shu Ma et al.ISCA 2023 · 15 citations
