Beehive: A Scalable Disaggregated Memory Runtime Exploiting Asynchrony of Multithreaded Programs
Quanxi Li, Hong Huang, Ying Liu, Yanwen Xia, Jie Zhang, Mosong Zhou, Xiaobing Feng, Huimin Cui, Quan Chen, Yizhou Shan, Chenxi Wang
Abstract
The Microsecond (µs)-scale I/O fabrics raise a tension between the programming productivity and performance, especially in disaggregated memory systems. The multithreaded synchronous programming model is popular in developing memory-disaggregated applications due to its intuitive program logic. However, our key insight is that although thread switching can effectively mitigate µs-scale latency, it leads to poor data locality and non-trivial scheduling overhead, leaving significant opportunities to improve the performance further. This paper proposes a memory-disaggregated framework, Beehive, which improves the remote access throughput by exploiting the asynchrony within each thread. To improve the programming usability, Beehive allows the programmers to develop applications in the conventional multithreaded synchronous model and automatically transforms the code into pararoutine (a newly proposed computation and scheduling unit) based asynchronous code via the Rust compiler. Beehive outperforms the state-of-the-art memory-disaggregated frameworks, i.e., Fastswap, Hermit, and AIFM, by 4.26×, 3.05×, and 1.58× on average.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Efficient and Flexible Datapaths for Fine-Grained Rack-Scale Interconnects with Elastic QPChenxingyu Zhao, Yibo Wu, Hongtao Zhang, Jaehong Min et al.SIGCOMM 2026 · 1 citation
- Shaving the Peaks: Taming Tail Latency for Managed Workloads via Disaggregated Garbage CollectionHongtao Lyu, Yuhan Li, Mingyu WuOSDI 2026
- Harvesting Sub-Microsecond CXL Memory Stalls with LiteSwitchNanqinqin Li, Yuhong Zhong, Asaf Cidon, Michael J. FreedmanOSDI 2026
- Efficient, Scalable, and Fair Locking on Disaggregated Memory with Decentralized CoordinationHanze Zhang, Ke Cheng, Rong Chen, Xingda Wei et al.VLDB 2026
Builds on23
- Pond: CXL-Based Memory Pooling Systems for Cloud PlatformsHuaicheng Li, Daniel S. Berger, Lisa Hsu, Daniel Ernst et al.ASPLOS 2023 · 328 citations
- AIFM: High-Performance, Application-Integrated Far MemoryZhenyuan Ruan, Malte Schwarzkopf, Marcos K. Aguilera, Adam BelayOSDI 2020 · 224 citations
- Effectively Prefetching Remote Memory with LeapHasan Al Maruf, Mosharaf ChowdhuryUSENIX ATC 2020 · 186 citations
- Can far memory improve job throughput?Emmanuel Amaro, Christopher Branner-Augmon, Zhihong Luo, Amy Ousterhout et al.EuroSys 2020 · 163 citations
- The CacheLib Caching Engine: Design and Experiences at ScaleBenjamin Berg, Daniel S. Berger, Sara McAllister, Isaac Grosof et al.OSDI 2020 · 145 citations
Related papers
- Cowbird: Freeing CPUs to Compute by Offloading the Disaggregation of MemoryXinyi Chen, Liangcheng Yu, Vincent Liu, Qizhen ZhangSIGCOMM 2023 · 15 citations
- pulse: Accelerating Distributed Pointer-Traversals on Disaggregated MemoryYupeng Tang, Seung-Seob Lee, Abhishek Bhattacharjee, Anurag KhandelwalASPLOS 2025 · 4 citations
- Adios to Busy-Waiting for Microsecond-scale Memory DisaggregationWonsup Yoon, Jisu Ok, Sue Moon, Youngjin KwonEuroSys 2025 · 3 citations
- CIDER: Boosting Memory-Disaggregated Key-Value Stores with Pessimistic SynchronizationYuxuan Du, Xuchuan Luo, Xin Wang, Yangfan Zhou et al.VLDB 2026
- Scaling Up Memory Disaggregated Applications with SMARTFeng Ren, Mingxing Zhang, Kang Chen, Huaxia Xia et al.ASPLOS 2024 · 16 citations
