MOST: Model-Based Compression with Outlier Storage for Time Series Data
Zehai Yang, Shimin Chen
Abstract
Time series data are used in a wide variety of applications. The explosive growth of the amount of time series data poses a significant challenge in efficient data storage and query processing. Unfortunately, existing compression techniques either show only low to medium compression ratio on time series data, or incur significant decompression overhead during query processing.
We propose a novel compression technique, MOST (Model-based compression with Outlier STorage) for time series data. As measurement values often change smoothly in a period of time, we divide a time series into segments of smooth changes, then compute a linear model for each segment. Since tiny errors are often acceptable in analysis tasks, we omit data points whose computed values are within a pre-specified error threshold from the actual values, thereby effectively reducing the data size. Outliers are rare but important for many applications, and therefore we store outliers explicitly. Moreover, for processing MOST compressed data, we propose a segment-outlier dual-mode query engine that computes segments as a whole as much as possible, and build a prototype MostDB. Experimental results on real-world data sets show that MOST achieves 9.45-15.04x compression ratios. Compared to existing time series databases, MostDB achieves up to 11.68x speedups for common queries from the IoTDB Benchmark.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4a3d8d6c-3e23-43bf-8f8e-881f36a4885dCited by top-tier papers4
- Beyond Compression: A Comprehensive Evaluation of Lossless Floating-Point CompressionKaisei Hishida, Chunwei Liu, John Paparrizos, Aaron J. ElmoreVLDB 2025 · 8 citations
- Serf: Streaming Error-Bounded Floating-Point CompressionRuiyuan Li, Zechao Chen, Ruyun Lu, Xiaolong Xu et al.SIGMOD 2025 · 7 citations
- Improving Time Series Data Compression in Apache IoTDBYuxin Tang, Feng Zhang, Jiawei Guan, Yuan Tian et al.VLDB 2025 · 2 citations
- DeXOR: Enabling XOR in Decimal Space for Streaming Lossless Compression of Floating-point DataChuanyi Lv, Huan Li, Dingyu Yang, Zhonele Xie et al.VLDB 2026
Builds on9
- Approximate Analytics System over Compressed Time Series with Tight Deterministic Error GuaranteesChunbin Lin, Etienne Boursier, Yannis PapakonstantinouVLDB 2020 · 505 citations
- ALEX: An Updatable Adaptive Learned IndexJialin Ding, Umar Farooq Minhas, Jia Yu, Chi Wang et al.SIGMOD 2020 · 274 citations
- Optimizing Error-Bounded Lossy Compression for Scientific Data by Dynamic Spline InterpolationKai Zhao, Sheng Di, Maxim Dmitriev, Thierry-Laurent D. Tonellot et al.ICDE 2021 · 151 citations
- SAND: Streaming Subsequence Anomaly DetectionPaul Boniol, John Paparrizos, Themis Palpanas, Michael J. FranklinVLDB 2021 · 128 citations
- Decomposed Bounded Floats for Fast Compression and QueriesChunwei Liu, Hao Jiang, John Paparrizos, Aaron J. ElmoreVLDB 2021 · 65 citations
Related papers
- REGER: Reordering Time Series Data for Regression EncodingJinzhao Xiao, Wendi He, Shaoxu Song, Xiangdong Huang et al.ICDE 2024 · 1 citation
- Time Series Representation for Visualization in Apache IoTDBLei Rui, Xiangdong Huang, Shaoxu Song, Yuyuan Kang et al.SIGMOD 2024 · 7 citations
- Scalable Model-Based Management of Correlated Dimensional Time Series in ModelarDB+Søren Kejser Jensen, Torben Bach Pedersen, Christian ThomsenICDE 2021 · 22 citations
- Distance-based Outlier Query Optimization in Apache IoTDBYunxiang Su, Shaoxu Song, Xiangdong Huang, Chen Wang et al.VLDB 2024 · 2 citations
- OneRoundSTL: In-Database Seasonal-Trend DecompositionZijie Chen, Shaoxu Song, Jianmin WangICDE 2025 · 1 citation
