Compressing High-Frequency Time Series Through Multiple Models and Stealing From Residuals
Abduvoris Abduvakhobov, Søren Kejser Jensen, Christian Thomsen, Torben Bach Pedersen
Abstract
Wind turbines are equipped with high-quality sensors that generate vast volumes of high-frequency time series. The time series are ingested on the edge and transferred to the cloud for later analytics. This process is complicated by challenges like low network bandwidth and high cloud storage costs. ModelarDB was proposed as a solution to efficiently manage time series across the entire pipeline by using so-called models for lossless or error-bounded lossy compression of time series. However, ModelarDB’s compression can be further improved through: 1) avoiding models that only represent few values by storing residuals (i.e., values that models fail to compress) explicitly with them; 2) exploiting error bounds even more through preprocessing; and 3) timestamp compression specialized for regular and irregular time series. We propose the multi-model compression method Fauna which uses 1) the novel model fitting method Platypus; 2) PMC and Swing for compressing values and; 3) the novel Macaque for compressing residuals and timestamps. Platypus is a model fitting method that uses different models for specialized compression of values and residuals. We then evaluate state-of-the-art lossless compression methods for 32-bit floats and propose preprocessing methods to add support for error-bounded compression. We present Macaque that includes MacaqueV and MacaqueTS. MacaqueV modifies Facebook Gorilla’s lossless compression method for 32-bit floats (GorillaV) and combines it with our novel preprocessing methods to now also enable error-bounded lossy compression. MacaqueTS is a lossless compression method for timestamps. Using only Platypus reduces ModelarDB’s storage use by up to 1.8x and significantly simplifies using the system. While also up to 7x better for lossless compression, ModelarDB with Fauna uses up to 2.5x less storage than ModelarDB and up to 14.5x, 7.2x, 17.5x and 14.2x less storage than ClickHouse, Apache IoTDB, Apache Parquet and TimescaleDB, respectively, with a realistic 1% error bound.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- Scalable Model-Based Management of Massive High Frequency Wind Turbine Data with ModelarDBAbduvoris Abduvakhobov, Søren Kejser Jensen, Torben Bach Pedersen, Christian ThomsenVLDB 2024 · 2 citations
- Scalable Model-Based Management of Correlated Dimensional Time Series in ModelarDB+Søren Kejser Jensen, Torben Bach Pedersen, Christian ThomsenICDE 2021 · 22 citations
- MOST: Model-Based Compression with Outlier Storage for Time Series DataZehai Yang, Shimin ChenSIGMOD 2024 · 9 citations
- REGER: Reordering Time Series Data for Regression EncodingJinzhao Xiao, Wendi He, Shaoxu Song, Xiangdong Huang et al.ICDE 2024 · 1 citation
- Kangaroo: Efficient Lossless Floating-Point Compression via Dynamic Reference SelectionShuo Li, Xiaochun Yang, Chunhui Shen, Yutong Han et al.SIGMOD 2026
