Time Series Representation for Visualization in Apache IoTDB
Lei Rui, Xiangdong Huang, Shaoxu Song, Yuyuan Kang, Chen Wang, Jianmin Wang
Abstract
When analyzing time series, often interactively, the analysts frequently demand to visualize instantly large-scale data stored in databases. M4 visualization selects the first, last, bottom and top data points in each pixel column to ensure pixel-perfectness of the two-color line chart visualization. While M4 already shows its preciseness of encasing time series in different scales into a fixed size of pixels, how to efficiently support M4 representation in a time series native database is still absent. It is worth noting that, to enable fast writes, the commodity time series database systems, such as Apache IoTDB or InfluxDB, employ LSM-Tree based storage. That is, a time series is segmented and stored in a number of chunks, with possibly out-of-order arrivals, i.e., disordered on timestamps. To implement M4, a natural idea is to merge online the chunks as a whole series, with costly merge sort on timestamps, and then perform M4 representation as in relational databases. In this study, we propose a novel chunk merge free approach called M4-LSM to accelerate M4 representation and visualization. In particular, we utilize the metadata of chunks to prune and avoid the costly merging of any chunk. Moreover, intra-chunk indexing and pruning are enabled for efficiently accessing the representation points, referring to the special properties of time series. Remarkably, the time series database native operator M4-LSM has been implemented in Apache IoTDB, an open-source time series database, and deployed in companies across various industries. In the experiments over real-world datasets, the proposed M4-LSM operator demonstrates high efficiency without sacrificing preciseness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aebb976a-9740-46f2-acca-e154d6665e25Cited by top-tier papers1
Ask how each one uses itBuilds on3
- KD-Box: Line-segment-based KD-tree for Interactive Exploration of Large-scale Time-Series DataYue Zhao, Yunhai Wang, Jian Zhang, Chi-Wing Fu et al.IEEE VIS 2021 · 45 citations
- Time Series Data Encoding for Efficient Storage: A Comparative Analysis in Apache IoTDBJinzhao Xiao, Yuxiang Huang, Changyu Hu, Shaoxu Song et al.VLDB 2022 · 37 citations
- On Repairing Timestamps for Regular Interval Time SeriesChenguang Fang, Shaoxu Song, Yinan MeiVLDB 2022 · 18 citations
Related papers
- In-Database Time Series ClusteringYunxiang Su, Kenny Ye Liang, Shaoxu SongSIGMOD 2025 · 3 citations
- Learning Autoregressive Model in LSM-Tree based StoreYunxiang Su, Wenxuan Ma, Shaoxu SongKDD 2023 · 4 citations
- On Reducing Space Amplification with Multi-Column Compaction in Apache IoTDBChenguang Fang, Zijie Chen, Shaoxu Song, Xiangdong Huang et al.VLDB 2024 · 1 citation
- Distance-based Outlier Query Optimization in Apache IoTDBYunxiang Su, Shaoxu Song, Xiangdong Huang, Chen Wang et al.VLDB 2024 · 2 citations
- OM3: An Ordered Multi-level Min-Max Representation for Interactive Progressive Visualization of Time SeriesYunhai Wang, Yuchun Wang, Xin Chen, Yue Zhao et al.SIGMOD 2023 · 8 citations
