Lune

USENIX ATC2020Top-tier venue

Peregreen - modular database for efficient storage of historical time series in cloud environments

Alexander A. Visheratin, Alexey Struckov, Semen Yufa, Alexey Muratov, Denis A. Nasonov, Nikolay Butakov, Yury Kuznetsov, Michael May

2020Year
20Citations
5Top-tier citations

Abstract

The rapid development of scientific and industrial areas, which rely on time series data processing, raises the demand for storage that would be able to process tens and hundreds of terabytes of data efficiently. And by efficiency, one should understand not only the speed of data processing operations execution but also the volume of the data stored and operational costs when deploying the storage in a production environment such as cloud.

In this paper, we propose a concept for storing and indexing numeric time series that allows creating compact data representations optimized for cloud storages and perform typical operations -uploading, extracting, sampling, statistical aggregations, and transformations -at high speed. Our modular database that implements the proposed approach -Peregreen -can achieve a throughput of 3 million entries per second for uploading and 48 million entries per second for extraction in Amazon EC2 while having only Amazon S3 as storage backend for all the data.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 53629ef0-4eec-43b2-96bf-fb85419c6f57

Cited by top-tier papers5

Ask how each one uses it

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines