ALP: Adaptive Lossless floating-Point Compression
Azim Afroozeh, Leonardo Kuffó, Peter Boncz
摘要
IEEE 754 doubles do not exactly represent most real values, introducing rounding errors in computations and [de]serialization to text. These rounding errors inhibit the use of existing lightweight compression schemes such as Delta and Frame Of Reference (FOR), but recently new schemes were proposed: Gorilla, Chimp128, Pseu-doDecimals (PDE), Elf and Patas. However, their compression ratios are not better than those of general-purpose compressors such as Zstd; while [de]compression is much slower than Delta and FOR. We propose and evaluate ALP, that significantly improves these previous schemes in both speed and compression ratio (Figure 1 ). We created ALP after carefully studying the datasets used to evaluate the previous schemes. To obtain speed, ALP is designed to fit vectorized execution. This turned out to be key for also improving the compression ratio, as we found in-vector commonalities to create compression opportunities. ALP is an adaptive scheme that uses a strongly enhanced version of PseudoDecimals [31] to losslessly encode doubles as integers if they originated as decimals, and otherwise uses vectorized compression of the doubles' front bits. Its high speeds stem from our implementation in scalar code that auto-vectorizes, using building blocks provided by our Fast-Lanes library [6] , and an efficient two-stage compression algorithm that first samples row-groups and then vectors.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- The FastLanes File FormatAzim Afroozeh, Peter BonczVLDB 2025 · 被引用 9 次
- Beyond Compression: A Comprehensive Evaluation of Lossless Floating-Point CompressionKaisei Hishida, Chunwei Liu, John Paparrizos, Aaron J. ElmoreVLDB 2025 · 被引用 8 次
- Serf: Streaming Error-Bounded Floating-Point CompressionRuiyuan Li, Zechao Chen, Ruyun Lu, Xiaolong Xu 等SIGMOD 2025 · 被引用 7 次
- PDX: A Data Layout for Vector Similarity SearchLeonardo Kuffó, Elena Krippner, Peter BonczSIGMOD 2025 · 被引用 6 次
- Learned Compression of Nonlinear Time Series with Random AccessAndrea Guerra, Giorgio Vinciguerra, Antonio Boffa, Paolo FerraginaICDE 2025 · 被引用 5 次
它引用的顶会 Paper5
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Chimp: Efficient Lossless Floating Point Compression for Time Series DatabasesPanagiotis Liakos, Katia Papakonstantinopoulou, Yannis KotidisVLDB 2022 · 被引用 76 次
- BtrBlocks: Efficient Columnar Compression for Data LakesMaximilian Kuschewski, David Sauerwein, Adnan Alhomssi, Viktor LeisSIGMOD 2023 · 被引用 47 次
- Elf: Erasing-based Lossless Floating-Point CompressionRuiyuan Li, Zheng Li, Yi Wu, Chao Chen 等VLDB 2023 · 被引用 44 次
- The FastLanes Compression Layout: Decoding >100 Billion Integers per Second with Scalar CodeAzim Afroozeh, Peter BonczVLDB 2023 · 被引用 44 次
相关 Paper
- Camel: Efficient Compression of Floating-Point Time SeriesYuanyuan Yao, Lu Chen, Ziquan Fang, Yunjun Gao 等SIGMOD 2025 · 被引用 4 次
- Efficient Lossless Compression of Scientific Floating-Point Data on CPUs and GPUsNoushin Azami, Alex Fallin, Martin BurtscherASPLOS 2025 · 被引用 18 次
- Everything You Always Wanted to Know About Storage Compressibility of Pre-Trained ML Models but Were Afraid to AskZhaoyuan Su, Ammar Ahmed, Zirui Wang, Ali Anwar 等VLDB 2024
- DeXOR: Enabling XOR in Decimal Space for Streaming Lossless Compression of Floating-point DataChuanyi Lv, Huan Li, Dingyu Yang, Zhonele Xie 等VLDB 2026
- MANS: Efficient and Portable ANS Encoding for Multi-Byte Integer Data on CPUs and GPUsWenjing Huang, Jinwu Yang, Shengquan Yin, Haoxu Li 等SC 2025 · 被引用 3 次
