Two-Level Data Compression using Machine Learning in Time Series Database
Xinyang Yu, Yanqing Peng, Feifei Li, Sheng Wang, Xiaowei Shen, Huijun Mai, Yue Xie
Abstract
The explosion of time series advances the development of time series databases. To reduce storage overhead in these systems, data compression is widely adopted. Most existing compression algorithms utilize the overall characteristics of the entire time series to achieve high compression ratio, but ignore local contexts around individual points. In this way, they are effective for certain data patterns, and may suffer inherent pattern changes in real-world time series. It is therefore strongly desired to have a compression method that can always achieve high compression ratio in the existence of pattern diversity.
In this paper, we propose a two-level compression model that selects a proper compression scheme for each individual point, so that diverse patterns can be captured at a fine granularity. Based on this model, we design and implement AMMMO framework, where a set of control parameters is defined to distill and categorize data patterns. At the top level, we evaluate each sub-sequence to fill in these parameters, generating a set of compression scheme candidates (i.e., major mode selection). At the bottom level, we choose the best scheme from these candidates for each data point respectively (i.e., sub-mode selection). To effectively handle diverse data patterns, we introduce a reinforcement learning based approach to learn parameter values automatically. Our experimental evaluation shows that our approach improves compression ratio by up to 120% (with an average of 50%), compared to other time-series compression methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6d14c543-25bd-437c-9757-fb0e6d966fa9Cited by top-tier papers12
- Optimizing Error-Bounded Lossy Compression for Scientific Data by Dynamic Spline InterpolationKai Zhao, Sheng Di, Maxim Dmitriev, Thierry-Laurent D. Tonellot et al.ICDE 2021 · 151 citations
- TRACE: A Fast Transformer-based General-Purpose Lossless CompressorYu Mao, Yufei Cui, Tei-Wei Kuo, Chun Jason XueWWW 2022 · 60 citations
- Elf: Erasing-based Lossless Floating-Point CompressionRuiyuan Li, Zheng Li, Yi Wu, Chao Chen et al.VLDB 2023 · 44 citations
- Time Series Data Encoding for Efficient Storage: A Comparative Analysis in Apache IoTDBJinzhao Xiao, Yuxiang Huang, Changyu Hu, Shaoxu Song et al.VLDB 2022 · 37 citations
- High-performance Effective Scientific Error-bounded Lossy Compression with Auto-tuned Multi-component InterpolationJinyang Liu, Sheng Di, Kai Zhao, Xin Liang et al.SIGMOD 2024 · 29 citations
Related papers
- MOST: Model-Based Compression with Outlier Storage for Time Series DataZehai Yang, Shimin ChenSIGMOD 2024 · 9 citations
- TVStore: Automatically Bounding Time Series Storage via Time-Varying CompressionYanzhe An, Yue Su, Yuqing Zhu, Jianmin WangFAST 2022 · 10 citations
- FLEA: Frequency-based Lossless Encoding Algorithm for Periodic Time SeriesTianrui Xia, Jinzhao Xiao, Shaoxu SongSIGMOD 2026
- Camel: Efficient Compression of Floating-Point Time SeriesYuanyuan Yao, Lu Chen, Ziquan Fang, Yunjun Gao et al.SIGMOD 2025 · 4 citations
- LeCo: Lightweight Compression via Learning Serial CorrelationsYihao Liu, Xinyu Zeng, Huanchen ZhangSIGMOD 2024 · 17 citations
