Peregreen - modular database for efficient storage of historical time series in cloud environments
Alexander A. Visheratin, Alexey Struckov, Semen Yufa, Alexey Muratov, Denis A. Nasonov, Nikolay Butakov, Yury Kuznetsov, Michael May
Abstract
The rapid development of scientific and industrial areas, which rely on time series data processing, raises the demand for storage that would be able to process tens and hundreds of terabytes of data efficiently. And by efficiency, one should understand not only the speed of data processing operations execution but also the volume of the data stored and operational costs when deploying the storage in a production environment such as cloud.
In this paper, we propose a concept for storing and indexing numeric time series that allows creating compact data representations optimized for cloud storages and perform typical operations -uploading, extracting, sampling, statistical aggregations, and transformations -at high speed. Our modular database that implements the proposed approach -Peregreen -can achieve a throughput of 3 million entries per second for uploading and 48 million entries per second for extraction in Amazon EC2 while having only Amazon S3 as storage backend for all the data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 53629ef0-4eec-43b2-96bf-fb85419c6f57Cited by top-tier papers5
- Heracles: An Efficient Storage Model And Data Flushing For Performance Monitoring TimeseriesZhiqi Wang, Jin Xue, Zili ShaoVLDB 2021 · 14 citations
- TSCache: An Efficient Flash-based Caching Scheme for Time-series Data WorkloadsJian Liu, Kefei Wang, Feng ChenVLDB 2021 · 12 citations
- TVStore: Automatically Bounding Time Series Storage via Time-Varying CompressionYanzhe An, Yue Su, Yuqing Zhu, Jianmin WangFAST 2022 · 10 citations
- ADAMAS: Adaptive Domain-Aware Performance Anomaly Detection in Cloud Service SystemsWenwei Gu, Jiazhen Gu, Jinyang Liu, Zhuangbin Chen et al.ICSE 2025 · 4 citations
- An Efficient Cloud Storage Model with Compacted Metadata Management for Performance Monitoring Timeseries SystemsKai Zhang, Tianyu Wang, Zili ShaoFAST 2026
Related papers
- MOST: Model-Based Compression with Outlier Storage for Time Series DataZehai Yang, Shimin ChenSIGMOD 2024 · 9 citations
- Time Series Data Encoding for Efficient Storage: A Comparative Analysis in Apache IoTDBJinzhao Xiao, Yuxiang Huang, Changyu Hu, Shaoxu Song et al.VLDB 2022 · 37 citations
- TimeUnion: An Efficient Architecture with Unified Data Model for Timeseries Management Systems on Hybrid Cloud StorageZhiqi Wang, Zili ShaoSIGMOD 2022 · 10 citations
- STsCache: An Efficient Semantic Caching Scheme for Time-series Data Workloads Based on Hybrid StorageTao Kong, Hui Li, Yuxuan Zhao, Liping Li et al.VLDB 2025
- ForestTI: A Scalable Inverted-Index-Oriented Timeseries Management System with Flexible Memory EfficiencyZhiqi Wang, Zili ShaoSIGMOD 2023 · 2 citations
