STsCache: An Efficient Semantic Caching Scheme for Time-series Data Workloads Based on Hybrid Storage
Tao Kong, Hui Li, Yuxuan Zhao, Liping Li, Xiyue Gao, Qilong Wu, Jiangtao Cui
Abstract
Due to the increasing demand for extreme-scale time-series data workloads in data centers, it is required to build a high-performance semantic caching system that leverages the semantics and results of historical queries to answer time-series queries. Existing caching solutions either ignore the semantics of queries, offering suboptimal performance, or focus only on specific scenarios, providing small-capacity, limited functionality. In this paper, we summarize the query patterns of time-series data workload and propose the definition of semantic time-series caching for the first time. Accordingly, we present a semantic time-series caching system, STsCache, based on a hybrid storage model with memory and NVMe SSD. We propose a series of optimized strategies, such as slab-based semantic data management, semantic index, semantic value-driven batch eviction, time-aware deduplication insertion, and lazy compaction. We implemented and evaluated STsCache via benchmarks and production environments. STsCache can increase throughput of popular time-series databases (InfluxDB, TimescaleDB) by 4.8–10.8X and reduce latency by 79.9%-93.5%. Compared with the latest time-series caching schemes (TSCache, BSCache), STsCache can increase throughput by 1.5–4.5X, reduce latency by 59.4%-81.9%, and increase hit ratios by 22.5%-82.4%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 54a4d3cd-13d2-4e89-ba11-de5fd8ef46caBuilds on12
- What Modern NVMe Storage Can Do, And How To Exploit It: High-Performance I/O for High-Performance Storage EnginesGabriel Haas, Viktor LeisVLDB 2023 · 83 citations
- TS-Benchmark: A Benchmark for Time Series DatabasesYuanzhe Hao, Xiongpai Qin, Yueguo Chen, Yaru Li et al.ICDE 2021 · 44 citations
- TreeLine: An Update-In-Place Key-Value Store for Modern StorageGeoffrey X. Yu, Markos Markakis, Andreas Kipf, Per-Åke Larson et al.VLDB 2023 · 36 citations
- TSM-Bench: Benchmarking Time Series Database Systems for Monitoring ApplicationsAbdelouahab Khelifati, Mourad Khayati, Anton Dignös, Djellel Eddine Difallah et al.VLDB 2023 · 24 citations
- Tri-Level Navigator: LLM-Empowered Tri-Level Learning for Time Series OOD GeneralizationChengtao Jian, Kai Yang, Yang JiaoNeurIPS 2024 · 19 citations
Related papers
- TSCache: An Efficient Flash-based Caching Scheme for Time-series Data WorkloadsJian Liu, Kefei Wang, Feng ChenVLDB 2021 · 12 citations
- Heracles: An Efficient Storage Model And Data Flushing For Performance Monitoring TimeseriesZhiqi Wang, Jin Xue, Zili ShaoVLDB 2021 · 14 citations
- A large scale analysis of hundreds of in-memory cache clusters at TwitterJuncheng Yang, Yao Yue, K. V. RashmiOSDI 2020 · 245 citations
- Visualization-aware Time Series Min-Max Caching with Error Bound GuaranteesStavros Maroulis, Vassilis Stamatopoulos, George Papastefanatos, Manolis TerrovitisVLDB 2024 · 8 citations
- Competitive Consistent Caching for TransactionsShuai An, Yang Cao, Wenyue ZhaoICDE 2022 · 1 citation
