Heracles: An Efficient Storage Model And Data Flushing For Performance Monitoring Timeseries
Zhiqi Wang, Jin Xue, Zili Shao
摘要
Performance-monitoring timeseries systems such as Prometheus and InfluxDB play a critical role in assuring reliability and operationality. These systems commonly adopt a column-oriented storage model, by which timeseries samples from different timeseries are separated, and all samples (with both numeric values and timestamps) in one timeseries are grouped into chunks and stored together. As a group of timeseries are often collected from the same source with the same timestamps, managing timestamps and metrics in a group manner provides more opportunities for query and insertion optimization but posts new challenges as well. Besides, for performance monitoring systems, to support better compression and efficient queries for most recent data that are most likely accessed by users, huge volumes of data are first cached in memory and then periodically flushed to disks. Periodic data flushing incurs high IO overhead, and simply discarding flushed data, which can still serve queries, not only is a waste but also brings huge memory reclamation cost. In this paper, we propose Heracles which integrates two techniques -( 1 ) a new storage model, which enables efficient queries on compressed data by utilizing the shared timestamp column to easily locate corresponding metric values; (2) a novel two-level epoch-based memory manager, which allows the system to gradually flush and reclaim in-memory data while unreclaimed data can still serve queries. Heracles is implemented as a standalone module that can be easily integrated into existing performance monitoring timeseries systems. We have implemented a fully functional prototype with Heracles based on Prometheus tsdb, a representative open-source performance monitoring system, and conducted extensive experiments with real and synthetic timeseries data. Experimental results show that, compared with Prometheus, Heracles can improve the insertion throughput by 171%, and reduce the query latency and space usage by 32% and 30%, respectively, on average. Besides, to compare with other state-of-the-art storage techniques, we have integrated LevelDB (for LSM-tree-based structure) and Parquet (for column stores) into Prometheus tsdb, respectively, and experimental results show Heracles outperform these two integrations. We have released the open-source code of Heracles for public access.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Grouping Time Series for Efficient Columnar StorageChenguang Fang, Shaoxu Song, Haoquan Guan, Xiangdong Huang 等SIGMOD 2023 · 被引用 10 次
- Improving Time Series Data Compression in Apache IoTDBYuxin Tang, Feng Zhang, Jiawei Guan, Yuan Tian 等VLDB 2025 · 被引用 2 次
- Approximation-First Timeseries Monitoring Query At ScaleZeying Zhu, Jonathan Chamberlain, Kenny Wu, David Starobinski 等VLDB 2025 · 被引用 2 次
- STsCache: An Efficient Semantic Caching Scheme for Time-series Data Workloads Based on Hybrid StorageTao Kong, Hui Li, Yuxuan Zhao, Liping Li 等VLDB 2025
它引用的顶会 Paper2
- An LSM-based Tuple Compaction Framework for Apache AsterixDBWail Y. Alkowaileet, Sattam Alsubaiee, Michael J. CareyVLDB 2020 · 被引用 23 次
- Peregreen - modular database for efficient storage of historical time series in cloud environmentsAlexander A. Visheratin, Alexey Struckov, Semen Yufa, Alexey Muratov 等USENIX ATC 2020 · 被引用 20 次
相关 Paper
- ForestTI: A Scalable Inverted-Index-Oriented Timeseries Management System with Flexible Memory EfficiencyZhiqi Wang, Zili ShaoSIGMOD 2023 · 被引用 2 次
- An Efficient Cloud Storage Model with Compacted Metadata Management for Performance Monitoring Timeseries SystemsKai Zhang, Tianyu Wang, Zili ShaoFAST 2026
- Time Series Data Encoding for Efficient Storage: A Comparative Analysis in Apache IoTDBJinzhao Xiao, Yuxiang Huang, Changyu Hu, Shaoxu Song 等VLDB 2022 · 被引用 37 次
- MOST: Model-Based Compression with Outlier Storage for Time Series DataZehai Yang, Shimin ChenSIGMOD 2024 · 被引用 9 次
- Deferred Flushing for Out-of-Order Arrivals in Apache IoTDBXiaojian Zhang, Zhiheng Liu, Shaoxu Song, Xiangdong Huang 等ICDE 2026
