An Efficient Cloud Storage Model with Compacted Metadata Management for Performance Monitoring Timeseries Systems
Kai Zhang, Tianyu Wang, Zili Shao
Abstract
Cloud-based performance monitoring timeseries systems are emerging due to their flexibility and pay-as-you-go capabilities. However, these systems encounter a major bottleneck in query performance, mainly attributed to the prolonged access latency of cloud storage and metadata redundancy of large number of timeseries. Thus, it is critical to optimize query performance within cloud environment and reduce metadata redundancy.
In this paper, we propose CloudTS, which is a novel timeseries data storage model with query optimization for cloud storage. CloudTS separately manages metadata and data, and introduces an efficient global metadata management for both space saving and query speedup. CloudTS also transparently supports the time-partitioned tag-based query model in performance monitoring timeseries systems. For metadata, a global tag dictionary is built to reduce metadata redundancy and a novel timeseries-tag mapping technique with a twodimension bitmap is designed so the mapping of timeseries and tags can be efficiently accomplished to support tag-based queries. For data, the compressed data chunks are put into objects by timeseries group. We have implemented a fully functional prototype of CloudTS and evaluated it with production timeseries data and synthetic workloads based on Amazon S3. In comparison, Cortex, a cloud-based timeseries system widely adopted by industries, and Apache Parquet and JSON Time Series, two representative cloud storage formats, are utilized in the evaluation. Experimental results show that CloudTS can improve query performance by 1.37× on average compared with Cortex, and outperforms Apache Parquet and JSON Time Series as well.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cf965594-4fd0-4c06-a19b-1e4c86efd427Builds on6
- Chimp: Efficient Lossless Floating Point Compression for Time Series DatabasesPanagiotis Liakos, Katia Papakonstantinopoulou, Yannis KotidisVLDB 2022 · 76 citations
- Time Series Data Encoding for Efficient Storage: A Comparative Analysis in Apache IoTDBJinzhao Xiao, Yuxiang Huang, Changyu Hu, Shaoxu Song et al.VLDB 2022 · 37 citations
- Scalable Model-Based Management of Correlated Dimensional Time Series in ModelarDB+Søren Kejser Jensen, Torben Bach Pedersen, Christian ThomsenICDE 2021 · 22 citations
- Peregreen - modular database for efficient storage of historical time series in cloud environmentsAlexander A. Visheratin, Alexey Struckov, Semen Yufa, Alexey Muratov et al.USENIX ATC 2020 · 20 citations
- TVStore: Automatically Bounding Time Series Storage via Time-Varying CompressionYanzhe An, Yue Su, Yuqing Zhu, Jianmin WangFAST 2022 · 10 citations
Related papers
- Heracles: An Efficient Storage Model And Data Flushing For Performance Monitoring TimeseriesZhiqi Wang, Jin Xue, Zili ShaoVLDB 2021 · 14 citations
- ForestTI: A Scalable Inverted-Index-Oriented Timeseries Management System with Flexible Memory EfficiencyZhiqi Wang, Zili ShaoSIGMOD 2023 · 2 citations
- TimeUnion: An Efficient Architecture with Unified Data Model for Timeseries Management Systems on Hybrid Cloud StorageZhiqi Wang, Zili ShaoSIGMOD 2022 · 10 citations
- A Spatio-Temporal Series Data Model with Efficient Indexing and Layout for Cloud-Based Trajectory Data ManagementYang Guo, Zhiqi Wang, Jin Xue, Zili ShaoICDE 2024 · 11 citations
- MOST: Model-Based Compression with Outlier Storage for Time Series DataZehai Yang, Shimin ChenSIGMOD 2024 · 9 citations
