TZ: Achieving High-Ratio Scientific Data Compression on GPUs with Global Data Decomposition
Zhuoxun Yang, Ruoyu Li, Amit N. Subrahmanya, Vishwas Rao, Sheng Di, Robert Underwood, Longtao Zhang, Jinyang Liu, Franck Cappello, Kai Zhao
Abstract
As high-performance computing shifts toward GPU-accelerated exascale systems, the exponential growth of scientific data poses severe challenges to both storage capacity and I/O bandwidth. While current GPU-based lossy compressors attempt to address this by porting CPU algorithms to the device, they rely heavily on block-wise spatial decomposition to fit GPU parallelism. This approach suffers from a fundamental locality barrier: by partitioning data into independent blocks, these methods fail to capture global correlations and fragment the unified data patterns required for effective coding, severely limiting compression ratios. In this paper, we propose TZ, a novel GPU-native error-bounded lossy compressor that breaks this ceiling by adopting global Tucker decomposition. By prioritizing global spectral energy compaction over local approximation, TZ naturally maximizes the compression potential for scientific datasets. To render this computationally intensive approach practical for high-throughput GPU workflows, we introduce a highly optimized adaptive randomized SVD engine. This design allows TZ to achieve the superior compression ratios of global spectral decomposition while maintaining competitive execution speeds. Furthermore, the global processing nature of TZ enables a unified quantization and coding scheme that eliminates block artifacts and metadata overhead. Evaluation on production-scale scientific datasets demonstrates that TZ achieves approximately 10 × higher compression ratios than state-of-the-art GPU compressors under the same error bound, while maintaining competitive, high-throughput performance.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get ff0265a8-c398-4070-a5fd-5a1e2f5b1b4bRelated papers
- FZ-GPU: A Fast and High-Ratio Lossy Compressor for Scientific Computing Applications on GPUsBoyuan Zhang, Jiannan Tian, Sheng Di, Xiaodong Yu et al.HPDC 2023 · 27 citations
- cuSZ-i: High-Ratio Scientific Lossy Compression on GPUs with Optimized Multi-Level InterpolationJinyang Liu, Jiannan Tian, Shixun Wu, Sheng Di et al.SC 2024 · 17 citations
- Boosting Scientific Error-Bounded Lossy Compression through Optimized Synergistic Lossy-Lossless OrchestrationShixun Wu, Jinwen Pan, Jinyang Liu, Jiannan Tian et al.SC 2025 · 6 citations
- Toward Scalable Tucker Decomposition: Skew-Aware Multi-Level Partitioning with GPU-Storage Co-ProcessingSeung Hyeon Song, Jihye Lee, Chanki Kim, Kang-Wook ChonICDE 2026
- cuSZp: An Ultra-fast GPU Error-bounded Lossy Compression Framework with Optimized End-to-End PerformanceYafan Huang, Sheng Di, Xiaodong Yu, Guanpeng Li et al.SC 2023 · 52 citations
