Para-ksm: Parallelized Memory Deduplication with Data Streaming Accelerator
Houxiang Ji, Minho Kim, Seonmu Oh, Daehoon Kim, Nam Sung Kim
Abstract
To tame the rapidly rising cost of memory in servers, hyperscalers have begun deploying memory deduplication features, such as Kernel Same-page Merging (ksm), for some of their services. Nonetheless, ksm incurs a datacenter tax significant enough to notably degrade performance of co-running applications, which hinders its wider and more aggressive deployment. Meanwhile, the server-class CPU has started to integrate various on-chip accelerators to effectively reduce datacenter taxes. One of such accelerators is Data Streaming Accelerator (DSA), which can offload the two most taxing functions of ksm, page comparison and checksum computation, from CPU. In this work, we demonstrate that ksm offloading these two functions to DSA (DSA-ksm) can reduce the performance degradation of co-running applications caused by ksm from 1.6-5.8× to 1.0-1.6×. However, we uncover that DSA-ksm, which naïvely replaces CPU-based functions with their DSA-based counterparts, yields significantly lower rates of memory deduplication than ksm due to the long latency of offloading these functions through on-chip PCIe. To address this shortcoming, we redesign ksm to exploit DSA's batching capability (Para-ksm). It facilitates a given function to operate on multiple pages per offload, rather than a single page as ksm does, thereby amortizing the long offloading latency. Compared to ksm, Para-ksm increases the amount of memory deduplication per CPU cycle used for ksm by 31-50% while decreasing the performance degradation to 1.3-2.7×.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bab73708-248c-4939-bfd1-76da0e5d1de7Builds on10
- TMO: transparent memory offloading in datacentersJohannes Weiner, Niket Agarwal, Dan Schatzberg, Leon Yang et al.ASPLOS 2022 · 103 citations
- LineFS: Efficient SmartNIC Offload of a Distributed File System with Pipeline ParallelismJongyul Kim, Insu Jang, Waleed Reda, Jaeseong Im et al.SOSP 2021 · 83 citations
- MEMTIS: Efficient Memory Tiering with Dynamic Page Classification and Page Size DeterminationTaehyung Lee, Sumit Kumar Monga, Changwoo Min, Young Ik EomSOSP 2023 · 67 citations
- FlexTOE: Flexible TCP Offload with Fine-Grained ParallelismRajath Shashidhara, Tim Stamler, Antoine Kaufmann, Simon PeterNSDI 2022 · 66 citations
- Memory deduplication for serverless computing with MedesDivyanshu Saxena, Tao Ji, Arjun Singhvi, Junaid Khalid et al.EuroSys 2022 · 54 citations
Related papers
- LightDSA: Enabling Efficient DSA Through Hardware-Aware Transparent OptimizationYuansen Wang, Teng Ma, Yuanhui Luo, Dongbiao He et al.EuroSys 2026
- A Quantitative Analysis and Guidelines of Data Streaming Accelerator in Modern Intel Xeon Scalable ProcessorsReese Kuper, Ipoom Jeong, Yifan Yuan, Ren Wang et al.ASPLOS 2024 · 24 citations
- BASK: Batch And SmartNIC-offloaded KSMChanshin Kwak, Jaehyeon Lee, Minkyu Jung, Changjun Lee et al.EuroSys 2026
- DSAssassin: Cross-VM Side-Channel Attacks by Exploiting Intel Data Streaming AcceleratorBen Chen, Kunlin Li, Shuwen Deng, Dongsheng Wang et al.HPCA 2026 · 1 citation
- STYX: Exploiting SmartNIC Capability to Reduce Datacenter Memory TaxHouxiang Ji, Mark Mansi, Yan Sun, Yifan Yuan et al.USENIX ATC 2023 · 22 citations
