An Efficient Cloud Storage Model with Compacted Metadata Management for Performance Monitoring Timeseries Systems
Kai Zhang, Tianyu Wang, Zili Shao
摘要
Cloud-based performance monitoring timeseries systems are emerging due to their flexibility and pay-as-you-go capabilities. However, these systems encounter a major bottleneck in query performance, mainly attributed to the prolonged access latency of cloud storage and metadata redundancy of large number of timeseries. Thus, it is critical to optimize query performance within cloud environment and reduce metadata redundancy.
In this paper, we propose CloudTS, which is a novel timeseries data storage model with query optimization for cloud storage. CloudTS separately manages metadata and data, and introduces an efficient global metadata management for both space saving and query speedup. CloudTS also transparently supports the time-partitioned tag-based query model in performance monitoring timeseries systems. For metadata, a global tag dictionary is built to reduce metadata redundancy and a novel timeseries-tag mapping technique with a twodimension bitmap is designed so the mapping of timeseries and tags can be efficiently accomplished to support tag-based queries. For data, the compressed data chunks are put into objects by timeseries group. We have implemented a fully functional prototype of CloudTS and evaluated it with production timeseries data and synthetic workloads based on Amazon S3. In comparison, Cortex, a cloud-based timeseries system widely adopted by industries, and Apache Parquet and JSON Time Series, two representative cloud storage formats, are utilized in the evaluation. Experimental results show that CloudTS can improve query performance by 1.37× on average compared with Cortex, and outperforms Apache Parquet and JSON Time Series as well.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Chimp: Efficient Lossless Floating Point Compression for Time Series DatabasesPanagiotis Liakos, Katia Papakonstantinopoulou, Yannis KotidisVLDB 2022 · 被引用 76 次
- Time Series Data Encoding for Efficient Storage: A Comparative Analysis in Apache IoTDBJinzhao Xiao, Yuxiang Huang, Changyu Hu, Shaoxu Song 等VLDB 2022 · 被引用 37 次
- Scalable Model-Based Management of Correlated Dimensional Time Series in ModelarDB+Søren Kejser Jensen, Torben Bach Pedersen, Christian ThomsenICDE 2021 · 被引用 22 次
- Peregreen - modular database for efficient storage of historical time series in cloud environmentsAlexander A. Visheratin, Alexey Struckov, Semen Yufa, Alexey Muratov 等USENIX ATC 2020 · 被引用 20 次
- TVStore: Automatically Bounding Time Series Storage via Time-Varying CompressionYanzhe An, Yue Su, Yuqing Zhu, Jianmin WangFAST 2022 · 被引用 10 次
相关 Paper
- Heracles: An Efficient Storage Model And Data Flushing For Performance Monitoring TimeseriesZhiqi Wang, Jin Xue, Zili ShaoVLDB 2021 · 被引用 14 次
- ForestTI: A Scalable Inverted-Index-Oriented Timeseries Management System with Flexible Memory EfficiencyZhiqi Wang, Zili ShaoSIGMOD 2023 · 被引用 2 次
- TimeUnion: An Efficient Architecture with Unified Data Model for Timeseries Management Systems on Hybrid Cloud StorageZhiqi Wang, Zili ShaoSIGMOD 2022 · 被引用 10 次
- A Spatio-Temporal Series Data Model with Efficient Indexing and Layout for Cloud-Based Trajectory Data ManagementYang Guo, Zhiqi Wang, Jin Xue, Zili ShaoICDE 2024 · 被引用 11 次
- MOST: Model-Based Compression with Outlier Storage for Time Series DataZehai Yang, Shimin ChenSIGMOD 2024 · 被引用 9 次
