Heracles: An Efficient Storage Model And Data Flushing For Performance Monitoring Timeseries
Zhiqi Wang, Jin Xue, Zili Shao
Abstract
Performance-monitoring timeseries systems such as Prometheus and InfluxDB play a critical role in assuring reliability and operationality. These systems commonly adopt a column-oriented storage model, by which timeseries samples from different timeseries are separated, and all samples (with both numeric values and timestamps) in one timeseries are grouped into chunks and stored together. As a group of timeseries are often collected from the same source with the same timestamps, managing timestamps and metrics in a group manner provides more opportunities for query and insertion optimization but posts new challenges as well. Besides, for performance monitoring systems, to support better compression and efficient queries for most recent data that are most likely accessed by users, huge volumes of data are first cached in memory and then periodically flushed to disks. Periodic data flushing incurs high IO overhead, and simply discarding flushed data, which can still serve queries, not only is a waste but also brings huge memory reclamation cost. In this paper, we propose Heracles which integrates two techniques -( 1 ) a new storage model, which enables efficient queries on compressed data by utilizing the shared timestamp column to easily locate corresponding metric values; (2) a novel two-level epoch-based memory manager, which allows the system to gradually flush and reclaim in-memory data while unreclaimed data can still serve queries. Heracles is implemented as a standalone module that can be easily integrated into existing performance monitoring timeseries systems. We have implemented a fully functional prototype with Heracles based on Prometheus tsdb, a representative open-source performance monitoring system, and conducted extensive experiments with real and synthetic timeseries data. Experimental results show that, compared with Prometheus, Heracles can improve the insertion throughput by 171%, and reduce the query latency and space usage by 32% and 30%, respectively, on average. Besides, to compare with other state-of-the-art storage techniques, we have integrated LevelDB (for LSM-tree-based structure) and Parquet (for column stores) into Prometheus tsdb, respectively, and experimental results show Heracles outperform these two integrations. We have released the open-source code of Heracles for public access.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext db89e14e-ef0b-4259-b116-db68d3897c36Cited by top-tier papers4
- Grouping Time Series for Efficient Columnar StorageChenguang Fang, Shaoxu Song, Haoquan Guan, Xiangdong Huang et al.SIGMOD 2023 · 10 citations
- Improving Time Series Data Compression in Apache IoTDBYuxin Tang, Feng Zhang, Jiawei Guan, Yuan Tian et al.VLDB 2025 · 2 citations
- Approximation-First Timeseries Monitoring Query At ScaleZeying Zhu, Jonathan Chamberlain, Kenny Wu, David Starobinski et al.VLDB 2025 · 2 citations
- STsCache: An Efficient Semantic Caching Scheme for Time-series Data Workloads Based on Hybrid StorageTao Kong, Hui Li, Yuxuan Zhao, Liping Li et al.VLDB 2025
Builds on2
- An LSM-based Tuple Compaction Framework for Apache AsterixDBWail Y. Alkowaileet, Sattam Alsubaiee, Michael J. CareyVLDB 2020 · 23 citations
- Peregreen - modular database for efficient storage of historical time series in cloud environmentsAlexander A. Visheratin, Alexey Struckov, Semen Yufa, Alexey Muratov et al.USENIX ATC 2020 · 20 citations
Related papers
- ForestTI: A Scalable Inverted-Index-Oriented Timeseries Management System with Flexible Memory EfficiencyZhiqi Wang, Zili ShaoSIGMOD 2023 · 2 citations
- An Efficient Cloud Storage Model with Compacted Metadata Management for Performance Monitoring Timeseries SystemsKai Zhang, Tianyu Wang, Zili ShaoFAST 2026
- Time Series Data Encoding for Efficient Storage: A Comparative Analysis in Apache IoTDBJinzhao Xiao, Yuxiang Huang, Changyu Hu, Shaoxu Song et al.VLDB 2022 · 37 citations
- MOST: Model-Based Compression with Outlier Storage for Time Series DataZehai Yang, Shimin ChenSIGMOD 2024 · 9 citations
- Deferred Flushing for Out-of-Order Arrivals in Apache IoTDBXiaojian Zhang, Zhiheng Liu, Shaoxu Song, Xiangdong Huang et al.ICDE 2026
