Grouping Time Series for Efficient Columnar Storage
Chenguang Fang, Shaoxu Song, Haoquan Guan, Xiangdong Huang, Chen Wang, Jianmin Wang
Abstract
Columnar storage is now an industry standard design in most open-source or commercial time series database products, making them HTAP systems. The time column of a time series serves as the key for identifying the other value column, namely single-column storage scheme. When multiple time series share a similar set of timestamps, very likely in a module of multiple sensors, it is natural to group them together, i.e., one time column identifies multiple value columns in a single-group storage scheme. While multiple value columns sharing the same time column reduce the space cost of repeating timestamps, it may introduce extra space cost for recording null values. The reason is that time series may not be exactly aligned on each timestamp, owing to missing values, distinct data collection frequencies, unsynchronized clocks and so on. The columngroups storage scheme is thus to divide columns into multiple groups, within which the value columns share the same time column. Unfortunately, the problem of finding the optimal column groups for the minimum space cost is highly challenging, NP-hard according to our analysis. Thereby, we propose a heuristic algorithm for automatically grouping time series for efficient columnar storage. The column groups storage has been deployed in Apache IoTDB, an open-source time series database. The extensive performance analysis, over real-world data from our industrial partners, demonstrates that the proposed column groups achieve near optimal storage, more concise than the storage of single-column or single-group schemes. Interestingly, both the flushing and querying time costs of column groups are comparable to those of single-column or singlegroup, i.e., without incurring extra time cost. CCS Concepts: • Information systems → Stream management.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8febee40-51d7-472d-bfc6-6a677e68af1dCited by top-tier papers2
- Improving Time Series Data Compression in Apache IoTDBYuxin Tang, Feng Zhang, Jiawei Guan, Yuan Tian et al.VLDB 2025 · 2 citations
- On Reducing Space Amplification with Multi-Column Compaction in Apache IoTDBChenguang Fang, Zijie Chen, Shaoxu Song, Xiangdong Huang et al.VLDB 2024 · 1 citation
Builds on4
- Columnar Storage and List-based Processing for Graph Database Management SystemsPranjal Gupta, Amine Mhedhbi, Semih SalihogluVLDB 2021 · 31 citations
- Jigsaw: A Data Storage and Query Processing Engine for Irregular Table PartitioningDonghe Kang, Ruochen Jiang, Spyros BlanasSIGMOD 2021 · 17 citations
- Heracles: An Efficient Storage Model And Data Flushing For Performance Monitoring TimeseriesZhiqi Wang, Jin Xue, Zili ShaoVLDB 2021 · 14 citations
- Representing Temporal Attributes for Schema MatchingYinan Mei, Shaoxu Song, Yunsu Lee, Jungho Park et al.KDD 2020 · 2 citations
Related papers
- Time Series Data Encoding for Efficient Storage: A Comparative Analysis in Apache IoTDBJinzhao Xiao, Yuxiang Huang, Changyu Hu, Shaoxu Song et al.VLDB 2022 · 37 citations
- In-Database Time Series ClusteringYunxiang Su, Kenny Ye Liang, Shaoxu SongSIGMOD 2025 · 3 citations
- Distance-based Outlier Query Optimization in Apache IoTDBYunxiang Su, Shaoxu Song, Xiangdong Huang, Chen Wang et al.VLDB 2024 · 2 citations
- Exploring SIMD Vectorization in Aggregation Pipelines for Encoded IoT DataRui Kang, Shaoxu Song, Jianmin WangICDE 2025
- REGER: Reordering Time Series Data for Regression EncodingJinzhao Xiao, Wendi He, Shaoxu Song, Xiangdong Huang et al.ICDE 2024 · 1 citation
