Enabling Efficient NVM-Based Text Analytics without Decompression
Xiaokun Fang, Feng Zhang, Junxiang Nong, Mingxing Zhang, Puyun Hu, Yunpeng Chai, Xiaoyong Du
摘要
Text analytics directly on compression (TADOC) is a promising technology designed for handling big data analytics. However, a substantial amount of DRAM is required for high performance, which limits its usage in many important scenarios where the capacity of DRAM is limited, such as memory-constrained systems. Non-volatile memory (NVM) is a novel storage technology that combines the advantage of reading per-formance and byte addressability of DRAM with the durability of traditional storage devices like SSD and HDD. Unfortunately, no research demonstrates how to use NVM to reduce DRAM utilization in compressed data analytics. In this paper, we propose N-TADOC, which substitutes DRAM with NVM while maintaining TADOC's analytics performance and space savings. Utilizing an NVM block device to reduce DRAM utilization presents two challenges, including poor data locality in traversing datasets and auxiliary data structure reconstruction on NVM. We develop novel designs to solve these challenges, including a pruning method with NVM pool management, bottom-up upper bound estimation, correspondent data structures, and persistence strategy at different levels of cost. Experimental results show that on four real-world datasets, N-TADOC achieves 2.04× performance speedup compared to the processing directly on the uncompressed data and 70.7% DRAM space saving compared to the original TADOC.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Assise: Performance and Availability via Client-local NVM in a Distributed File SystemThomas E. Anderson, Marco Canini, Jongyul Kim, Dejan Kostic 等OSDI 2020 · 被引用 71 次
- Twizzler: a Data-Centric OS for Non-Volatile MemoryDaniel Bittman, Peter Alvaro, Pankaj Mehra, Darrell D. E. Long 等USENIX ATC 2020 · 被引用 50 次
- PLIN: A Persistent Learned Index for Non-Volatile Memory with High Performance and Instant RecoveryZhou Zhang, Zhaole Chu, Peiquan Jin, Yongping Luo 等VLDB 2023 · 被引用 39 次
- G-TADOC: Enabling Efficient GPU-Based Text Analytics without DecompressionFeng Zhang, Zaifeng Pan, Yanliang Zhou, Jidong Zhai 等ICDE 2021 · 被引用 31 次
- CompressGraph: Efficient Parallel Graph Analytics with Rule-Based CompressionZheng Chen, Feng Zhang, Jiawei Guan, Jidong Zhai 等SIGMOD 2023 · 被引用 23 次
相关 Paper
- F-TADOC: FPGA-Based Text Analytics Directly on Compression with HLSYanliang Zhou, Feng Zhang, Tuo Lin, Yuanjie Huang 等ICDE 2024 · 被引用 1 次
- HeMem: Scalable Tiered Memory Management for Big Data Applications and Real NVMAmanda Raybuck, Tim Stamler, Wei Zhang, Mattan Erez 等SOSP 2021 · 被引用 93 次
- Zen: a High-Throughput Log-Free OLTP Engine for Non-Volatile Main MemoryGang Liu, Leying Chen, Shimin ChenVLDB 2021 · 被引用 31 次
- Reducing Bit Writes in Non-volatile Main Memory by Similarity-aware CompressionZhangyu Chen, Yu Hua, Pengfei Zuo, Yuanyuan Sun 等DAC 2020 · 被引用 7 次
- HOOP: Efficient Hardware-Assisted Out-of-Place Update for Non-Volatile MemoryMiao Cai, Chance C. Coats, Jian HuangISCA 2020 · 被引用 33 次
