SC2021Top-tier venue
ndzip-gpu: efficient lossless compression of scientific floating-point data on GPUs
Fabian Knorr, Peter Thoman, Thomas Fahringer
Abstract
Lossless data compression is a promising software approach for reducing the bandwidth requirements of scientific applications on accelerator clusters without introducing approximation errors. Suitable compressors must be able to effectively compact floatingpoint data while saturating the system interconnect to avoid introducing unnecessary latencies.
We present ndzip-gpu, a novel, highly-efficient GPU parallelization scheme for the block compressor ndzip, which has recently set a new milestone in CPU floating-point compression speeds.
Through the combination of intra-block parallelism and efficient memory access patterns, ndzip-gpu achieves high resource utilization in decorrelating multi-dimensional data via the Integer Lorenzo Transform. We further introduce a novel, efficient warp-cooperative primitive for vertical bit packing, providing a high-throughput data reduction and expansion step.
Using a representative set of scientific data, we compare the performance of ndzip-gpu against five other, existing GPU compressors. While observing that effectiveness of any compressor strongly depends on characteristics of the dataset, we demonstrate that ndzip-gpu offers the best average compression ratio for the examined data. On Nvidia Turing, Volta and Ampere hardware, it achieves the highest single-precision throughput by a significant margin while maintaining a favorable trade-off between data reduction and throughput in the double-precision case.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- FZ-GPU: A Fast and High-Ratio Lossy Compressor for Scientific Computing Applications on GPUsBoyuan Zhang, Jiannan Tian, Sheng Di, Xiaodong Yu et al.HPDC 2023 · 27 citations
- An Empirical Study on Low GPU Utilization of Deep Learning JobsYanjie Gao, Yichen He, Xinze Li, Bo Zhao et al.ICSE 2024 · 22 citations
- ENEC: A Lossless AI Model Compression Method Enabling Fast Inference on Ascend NPUsJinwu Yang, Jiaan Wu, Zedong Liu, Xinyang Ma et al.ISCA 2026 · 3 citations
- Reducing the GPU Memory Bottleneck with Lossless Compression for MLAditya K. Kamath, Arvind Krishnamurthy, Marco Canini, Simon PeterEuroSys 2026
- ThreadFuser: A SIMT Analysis Framework for MIMD ProgramsAhmad Alawneh, Ni Kang, Mahmoud Khairy, Timothy G. RogersMICRO 2024
Related papers
- Efficient Lossless Compression of Scientific Floating-Point Data on CPUs and GPUsNoushin Azami, Alex Fallin, Martin BurtscherASPLOS 2025 · 18 citations
- TZ: Achieving High-Ratio Scientific Data Compression on GPUs with Global Data DecompositionZhuoxun Yang, Ruoyu Li, Amit N. Subrahmanya, Vishwas Rao et al.HPDC 2026
- cuSZ-i: High-Ratio Scientific Lossy Compression on GPUs with Optimized Multi-Level InterpolationJinyang Liu, Jiannan Tian, Shixun Wu, Sheng Di et al.SC 2024 · 17 citations
- cuSZp: An Ultra-fast GPU Error-bounded Lossy Compression Framework with Optimized End-to-End PerformanceYafan Huang, Sheng Di, Xiaodong Yu, Guanpeng Li et al.SC 2023 · 52 citations
- PRISM: An Efficient GPU-Based Lossy Compression Framework for Progressive Data Retrieval with Multi-Level InterpolationBing Lu, Zedong Liu, Hairui Zhao, Dejun Luo et al.PPoPP 2026 · 2 citations
