Beehive: A Scalable Disaggregated Memory Runtime Exploiting Asynchrony of Multithreaded Programs
Quanxi Li, Hong Huang, Ying Liu, Yanwen Xia, Jie Zhang, Mosong Zhou, Xiaobing Feng, Huimin Cui, Quan Chen, Yizhou Shan, Chenxi Wang
摘要
The Microsecond (µs)-scale I/O fabrics raise a tension between the programming productivity and performance, especially in disaggregated memory systems. The multithreaded synchronous programming model is popular in developing memory-disaggregated applications due to its intuitive program logic. However, our key insight is that although thread switching can effectively mitigate µs-scale latency, it leads to poor data locality and non-trivial scheduling overhead, leaving significant opportunities to improve the performance further. This paper proposes a memory-disaggregated framework, Beehive, which improves the remote access throughput by exploiting the asynchrony within each thread. To improve the programming usability, Beehive allows the programmers to develop applications in the conventional multithreaded synchronous model and automatically transforms the code into pararoutine (a newly proposed computation and scheduling unit) based asynchronous code via the Rust compiler. Beehive outperforms the state-of-the-art memory-disaggregated frameworks, i.e., Fastswap, Hermit, and AIFM, by 4.26×, 3.05×, and 1.58× on average.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Efficient and Flexible Datapaths for Fine-Grained Rack-Scale Interconnects with Elastic QPChenxingyu Zhao, Yibo Wu, Hongtao Zhang, Jaehong Min 等SIGCOMM 2026 · 被引用 1 次
- Shaving the Peaks: Taming Tail Latency for Managed Workloads via Disaggregated Garbage CollectionHongtao Lyu, Yuhan Li, Mingyu WuOSDI 2026
- Harvesting Sub-Microsecond CXL Memory Stalls with LiteSwitchNanqinqin Li, Yuhong Zhong, Asaf Cidon, Michael J. FreedmanOSDI 2026
- Efficient, Scalable, and Fair Locking on Disaggregated Memory with Decentralized CoordinationHanze Zhang, Ke Cheng, Rong Chen, Xingda Wei 等VLDB 2026
它引用的顶会 Paper23
- Pond: CXL-Based Memory Pooling Systems for Cloud PlatformsHuaicheng Li, Daniel S. Berger, Lisa Hsu, Daniel Ernst 等ASPLOS 2023 · 被引用 328 次
- AIFM: High-Performance, Application-Integrated Far MemoryZhenyuan Ruan, Malte Schwarzkopf, Marcos K. Aguilera, Adam BelayOSDI 2020 · 被引用 224 次
- Effectively Prefetching Remote Memory with LeapHasan Al Maruf, Mosharaf ChowdhuryUSENIX ATC 2020 · 被引用 186 次
- Can far memory improve job throughput?Emmanuel Amaro, Christopher Branner-Augmon, Zhihong Luo, Amy Ousterhout 等EuroSys 2020 · 被引用 163 次
- The CacheLib Caching Engine: Design and Experiences at ScaleBenjamin Berg, Daniel S. Berger, Sara McAllister, Isaac Grosof 等OSDI 2020 · 被引用 145 次
相关 Paper
- Cowbird: Freeing CPUs to Compute by Offloading the Disaggregation of MemoryXinyi Chen, Liangcheng Yu, Vincent Liu, Qizhen ZhangSIGCOMM 2023 · 被引用 15 次
- pulse: Accelerating Distributed Pointer-Traversals on Disaggregated MemoryYupeng Tang, Seung-Seob Lee, Abhishek Bhattacharjee, Anurag KhandelwalASPLOS 2025 · 被引用 4 次
- Adios to Busy-Waiting for Microsecond-scale Memory DisaggregationWonsup Yoon, Jisu Ok, Sue Moon, Youngjin KwonEuroSys 2025 · 被引用 3 次
- CIDER: Boosting Memory-Disaggregated Key-Value Stores with Pessimistic SynchronizationYuxuan Du, Xuchuan Luo, Xin Wang, Yangfan Zhou 等VLDB 2026
- Scaling Up Memory Disaggregated Applications with SMARTFeng Ren, Mingxing Zhang, Kang Chen, Huaxia Xia 等ASPLOS 2024 · 被引用 16 次
