USENIX ATC2025顶会
Para-ksm: Parallelized Memory Deduplication with Data Streaming Accelerator
Houxiang Ji, Minho Kim, Seonmu Oh, Daehoon Kim, Nam Sung Kim
摘要
To tame the rapidly rising cost of memory in servers, hyperscalers have begun deploying memory deduplication features, such as Kernel Same-page Merging (ksm), for some of their services. Nonetheless, ksm incurs a datacenter tax significant enough to notably degrade performance of co-running applications, which hinders its wider and more aggressive deployment. Meanwhile, the server-class CPU has started to integrate various on-chip accelerators to effectively reduce datacenter taxes. One of such accelerators is Data Streaming Accelerator (DSA), which can offload the two most taxing functions of ksm, page comparison and checksum computation, from CPU. In this work, we demonstrate that ksm offloading these two functions to DSA (DSA-ksm) can reduce the performance degradation of co-running applications caused by ksm from 1.6-5.8× to 1.0-1.6×. However, we uncover that DSA-ksm, which naïvely replaces CPU-based functions with their DSA-based counterparts, yields significantly lower rates of memory deduplication than ksm due to the long latency of offloading these functions through on-chip PCIe. To address this shortcoming, we redesign ksm to exploit DSA's batching capability (Para-ksm). It facilitates a given function to operate on multiple pages per offload, rather than a single page as ksm does, thereby amortizing the long offloading latency. Compared to ksm, Para-ksm increases the amount of memory deduplication per CPU cycle used for ksm by 31-50% while decreasing the performance degradation to 1.3-2.7×.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- TMO: transparent memory offloading in datacentersJohannes Weiner, Niket Agarwal, Dan Schatzberg, Leon Yang 等ASPLOS 2022 · 被引用 103 次
- LineFS: Efficient SmartNIC Offload of a Distributed File System with Pipeline ParallelismJongyul Kim, Insu Jang, Waleed Reda, Jaeseong Im 等SOSP 2021 · 被引用 83 次
- MEMTIS: Efficient Memory Tiering with Dynamic Page Classification and Page Size DeterminationTaehyung Lee, Sumit Kumar Monga, Changwoo Min, Young Ik EomSOSP 2023 · 被引用 67 次
- FlexTOE: Flexible TCP Offload with Fine-Grained ParallelismRajath Shashidhara, Tim Stamler, Antoine Kaufmann, Simon PeterNSDI 2022 · 被引用 66 次
- Memory deduplication for serverless computing with MedesDivyanshu Saxena, Tao Ji, Arjun Singhvi, Junaid Khalid 等EuroSys 2022 · 被引用 54 次
相关 Paper
- LightDSA: Enabling Efficient DSA Through Hardware-Aware Transparent OptimizationYuansen Wang, Teng Ma, Yuanhui Luo, Dongbiao He 等EuroSys 2026
- A Quantitative Analysis and Guidelines of Data Streaming Accelerator in Modern Intel Xeon Scalable ProcessorsReese Kuper, Ipoom Jeong, Yifan Yuan, Ren Wang 等ASPLOS 2024 · 被引用 24 次
- BASK: Batch And SmartNIC-offloaded KSMChanshin Kwak, Jaehyeon Lee, Minkyu Jung, Changjun Lee 等EuroSys 2026
- DSAssassin: Cross-VM Side-Channel Attacks by Exploiting Intel Data Streaming AcceleratorBen Chen, Kunlin Li, Shuwen Deng, Dongsheng Wang 等HPCA 2026 · 被引用 1 次
- STYX: Exploiting SmartNIC Capability to Reduce Datacenter Memory TaxHouxiang Ji, Mark Mansi, Yan Sun, Yifan Yuan 等USENIX ATC 2023 · 被引用 22 次
