Revitalizing the Forgotten On-Chip DMA to Expedite Data Movement in NVM-based Storage Systems
Jingbo Su, Jiahao Li, Luofan Chen, Cheng Li, Kai Zhang, Liang Yang, Sam H. Noh, Yinlong Xu
摘要
Data-intensive applications executing on NVM-based storage systems experience serious bottlenecks when moving data between DRAM and NVM. We advocate for the use of the long-existing but recently neglected on-chip DMA to expedite data movement with three contributions. First, we explore new latency-oriented optimization directions, driven by a comprehensive DMA study, to design a high-performance DMA module, which significantly lowers the I/O size threshold to observe benefits. Second, we propose a new data movement engine, Fastmove, that coordinates the use of the DMA along with the CPU with judicious scheduling and load splitting such that the DMA's limitations are compensated, and the overall gains are maximized. Finally, with a general kernel-based design, simple APIs, and DAX file system integration, Fastmove allows applications to transparently exploit the DMA and its new features without code change. We run three data-intensive applications MySQL, Graph-Walker, and Filebench atop NOVA, ext4-DAX, and XFS-DAX, with standard benchmarks like TPC-C, and popular graph algorithms like PageRank. Across single-and multi-socket settings, compared to the conventional CPU-only NVM accesses, Fastmove introduces to TPC-C with MySQL 1.13-2.16× speedups of peak throughput, reduces the average latency by 17.7-60.8%, and saves 37.1-68.9% CPU usage spent in data movement. It also shortens the execution time of graph algorithms with GraphWalker by 39.7-53.4%, and introduces 1.12-1.27× throughput speedups for Filebench.
- This work was done at UNIST. † "SmartX" is also known as Beijing Zhiling Haina Technology Co., LTd.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Demystifying CXL Memory with Genuine CXL-Ready Systems and DevicesYan Sun, Yifan Yuan, Zeduo Yu, Reese Kuper 等MICRO 2023 · 被引用 133 次
- A Quantitative Analysis and Guidelines of Data Streaming Accelerator in Modern Intel Xeon Scalable ProcessorsReese Kuper, Ipoom Jeong, Yifan Yuan, Ren Wang 等ASPLOS 2024 · 被引用 24 次
它引用的顶会 Paper20
- An Empirical Guide to the Behavior and Use of Scalable Persistent MemoryJian Yang, Juno Kim, Morteza Hoseinzadeh, Joseph Izraelevitz 等FAST 2020 · 被引用 470 次
- HeMem: Scalable Tiered Memory Management for Big Data Applications and Real NVMAmanda Raybuck, Tim Stamler, Wei Zhang, Mattan Erez 等SOSP 2021 · 被引用 93 次
- Viper: An Efficient Hybrid PMem-DRAM Key-Value StoreLawrence Benson, Hendrik Makait, Tilmann RablVLDB 2021 · 被引用 86 次
- LineFS: Efficient SmartNIC Offload of a Distributed File System with Pipeline ParallelismJongyul Kim, Insu Jang, Waleed Reda, Jaeseong Im 等SOSP 2021 · 被引用 83 次
- Assise: Performance and Availability via Client-local NVM in a Distributed File SystemThomas E. Anderson, Marco Canini, Jongyul Kim, Dejan Kostic 等OSDI 2020 · 被引用 71 次
相关 Paper
- Exploring the Asynchrony of Slow Memory Filesystem with EasyIOBohong Zhu, Youmin Chen, Jiwu ShuEuroSys 2024 · 被引用 4 次
- Characterizing and Optimizing Remote Persistent Memory with RDMA and NVMXingda Wei, Xiating Xie, Rong Chen, Haibo Chen 等USENIX ATC 2021 · 被引用 48 次
- D-Shield: Enabling Processor-side Encryption and Integrity Verification for Secure NVMe DrivesMd Hafizul Islam Chowdhuryy, Myoungsoo Jung, Fan Yao, Amro AwadHPCA 2023 · 被引用 7 次
- Simurgh: a fully decentralized and secure NVMM user space file systemNafiseh Moti, Frederic Schimmelpfennig, Reza Salkhordeh, David Klopp 等SC 2021 · 被引用 11 次
- HAIL-DIMM: Host Access Interleaved with Near-Data Processing on DIMM-based Memory SystemMinkyu Lee, Sang-Seol Lee, Kyungho Kim, Eunchong Lee 等DAC 2024 · 被引用 2 次
