Learned Compression of Nonlinear Time Series with Random Access
Andrea Guerra, Giorgio Vinciguerra, Antonio Boffa, Paolo Ferragina
摘要
Time series play a crucial role in many fields, including finance, healthcare, industry, and environmental monitoring. The storage and retrieval of time series can be challenging due to their unstoppable growth. In fact, these applications often sacrifice precious historical data to make room for new data.
General-purpose compressors like Xz and Zstd can mitigate this problem with their good compression ratios, but they lack efficient random access on compressed data, thus preventing realtime analyses. Ad-hoc streaming solutions, instead, typically optimise only for compression and decompression speed, while giving up compression effectiveness and random access functionality. Furthermore, all these methods lack awareness of certain special regularities of time series, whose trends over time can often be described by some linear and nonlinear functions.
To address these issues, we introduce NeaTS, a randomlyaccessible compression scheme that approximates the time series with a sequence of nonlinear functions of different kinds and shapes, carefully selected and placed by a partitioning algorithm to minimise the space. The approximation residuals are bounded, which allows storing them in little space and thus recovering the original data losslessly, or simply discarding them to obtain a lossy time series representation with maximum error guarantees.
Our experiments show that NeaTS improves the compression ratio of the state-of-the-art lossy compressors that use linear or nonlinear functions (or both) by up to 14%. Compared to lossless compressors, NeaTS emerges as the only approach to date providing, simultaneously, compression ratios close to or better than the best existing compressors, a much faster decompression speed, and orders of magnitude more efficient random access, thus enabling the storage and real-time analysis of massive and ever-growing amounts of (historical) time series data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Learned Static Function Data StructuresStefan Hermann, Hans-Peter Lehmann, Giorgio Vinciguerra, Stefan WalzerVLDB 2026
- Eureka: Enabling Fine-Grained Access and Range Queries on Compressed Scientific Data via Data-Index Co-CompressionNing Yan, Sheng Di, Kai Zhao, Lipeng WanVLDB 2026
- LogFold: Compressing Logs with Structured Tokens and Hybrid EncodingShiwen Shan, Yintong Huo, Hongzhan Zhong, Zhining Wang 等ICSE 2026
- DeXOR: Enabling XOR in Decimal Space for Streaming Lossless Compression of Floating-point DataChuanyi Lv, Huan Li, Dingyu Yang, Zhonele Xie 等VLDB 2026
它引用的顶会 Paper15
- The PGM-index: a fully-dynamic compressed learned index with provable worst-case boundsPaolo Ferragina, Giorgio VinciguerraVLDB 2020 · 被引用 178 次
- Learned Index: A Comprehensive Experimental EvaluationZhaoyan Sun, Xuanhe Zhou, Guoliang LiVLDB 2023 · 被引用 87 次
- Chimp: Efficient Lossless Floating Point Compression for Time Series DatabasesPanagiotis Liakos, Katia Papakonstantinopoulou, Yannis KotidisVLDB 2022 · 被引用 76 次
- Are Updatable Learned Indexes Ready?Chaichon Wongkham, Baotong Lu, Chris Liu, Zhicong Zhong 等VLDB 2022 · 被引用 66 次
- Decomposed Bounded Floats for Fast Compression and QueriesChunwei Liu, Hao Jiang, John Paparrizos, Aaron J. ElmoreVLDB 2021 · 被引用 65 次
相关 Paper
- A Randomly Accessible Lossless Compression Scheme for Time-Series DataRasmus Vestergaard, Daniel E. Lucani, Qi ZhangINFOCOM 2020 · 被引用 29 次
- FLEA: Frequency-based Lossless Encoding Algorithm for Periodic Time SeriesTianrui Xia, Jinzhao Xiao, Shaoxu SongSIGMOD 2026
- ADT-FSE: A New Encoder for SZTao Lu, Yu Zhong, Zibin Sun, Xiang Chen 等SC 2023 · 被引用 10 次
- Camel: Efficient Compression of Floating-Point Time SeriesYuanyuan Yao, Lu Chen, Ziquan Fang, Yunjun Gao 等SIGMOD 2025 · 被引用 4 次
- STZ: A High Quality and High Speed Streaming Lossy Compression Framework for Scientific DataDaoce Wang, Pascal Grosset, Jesus Pulido, Jiannan Tian 等SC 2025 · 被引用 2 次
