ndzip-gpu: efficient lossless compression of scientific floating-point data on GPUs
Fabian Knorr, Peter Thoman, Thomas Fahringer
摘要
Lossless data compression is a promising software approach for reducing the bandwidth requirements of scientific applications on accelerator clusters without introducing approximation errors. Suitable compressors must be able to effectively compact floatingpoint data while saturating the system interconnect to avoid introducing unnecessary latencies.
We present ndzip-gpu, a novel, highly-efficient GPU parallelization scheme for the block compressor ndzip, which has recently set a new milestone in CPU floating-point compression speeds.
Through the combination of intra-block parallelism and efficient memory access patterns, ndzip-gpu achieves high resource utilization in decorrelating multi-dimensional data via the Integer Lorenzo Transform. We further introduce a novel, efficient warp-cooperative primitive for vertical bit packing, providing a high-throughput data reduction and expansion step.
Using a representative set of scientific data, we compare the performance of ndzip-gpu against five other, existing GPU compressors. While observing that effectiveness of any compressor strongly depends on characteristics of the dataset, we demonstrate that ndzip-gpu offers the best average compression ratio for the examined data. On Nvidia Turing, Volta and Ampere hardware, it achieves the highest single-precision throughput by a significant margin while maintaining a favorable trade-off between data reduction and throughput in the double-precision case.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- FZ-GPU: A Fast and High-Ratio Lossy Compressor for Scientific Computing Applications on GPUsBoyuan Zhang, Jiannan Tian, Sheng Di, Xiaodong Yu 等HPDC 2023 · 被引用 27 次
- An Empirical Study on Low GPU Utilization of Deep Learning JobsYanjie Gao, Yichen He, Xinze Li, Bo Zhao 等ICSE 2024 · 被引用 22 次
- ENEC: A Lossless AI Model Compression Method Enabling Fast Inference on Ascend NPUsJinwu Yang, Jiaan Wu, Zedong Liu, Xinyang Ma 等ISCA 2026 · 被引用 3 次
- Reducing the GPU Memory Bottleneck with Lossless Compression for MLAditya K. Kamath, Arvind Krishnamurthy, Marco Canini, Simon PeterEuroSys 2026
- ThreadFuser: A SIMT Analysis Framework for MIMD ProgramsAhmad Alawneh, Ni Kang, Mahmoud Khairy, Timothy G. RogersMICRO 2024
相关 Paper
- Efficient Lossless Compression of Scientific Floating-Point Data on CPUs and GPUsNoushin Azami, Alex Fallin, Martin BurtscherASPLOS 2025 · 被引用 18 次
- TZ: Achieving High-Ratio Scientific Data Compression on GPUs with Global Data DecompositionZhuoxun Yang, Ruoyu Li, Amit N. Subrahmanya, Vishwas Rao 等HPDC 2026
- cuSZ-i: High-Ratio Scientific Lossy Compression on GPUs with Optimized Multi-Level InterpolationJinyang Liu, Jiannan Tian, Shixun Wu, Sheng Di 等SC 2024 · 被引用 17 次
- cuSZp: An Ultra-fast GPU Error-bounded Lossy Compression Framework with Optimized End-to-End PerformanceYafan Huang, Sheng Di, Xiaodong Yu, Guanpeng Li 等SC 2023 · 被引用 52 次
- PRISM: An Efficient GPU-Based Lossy Compression Framework for Progressive Data Retrieval with Multi-Level InterpolationBing Lu, Zedong Liu, Hairui Zhao, Dejun Luo 等PPoPP 2026 · 被引用 2 次
