STsCache: An Efficient Semantic Caching Scheme for Time-series Data Workloads Based on Hybrid Storage
Tao Kong, Hui Li, Yuxuan Zhao, Liping Li, Xiyue Gao, Qilong Wu, Jiangtao Cui
摘要
Due to the increasing demand for extreme-scale time-series data workloads in data centers, it is required to build a high-performance semantic caching system that leverages the semantics and results of historical queries to answer time-series queries. Existing caching solutions either ignore the semantics of queries, offering suboptimal performance, or focus only on specific scenarios, providing small-capacity, limited functionality. In this paper, we summarize the query patterns of time-series data workload and propose the definition of semantic time-series caching for the first time. Accordingly, we present a semantic time-series caching system, STsCache, based on a hybrid storage model with memory and NVMe SSD. We propose a series of optimized strategies, such as slab-based semantic data management, semantic index, semantic value-driven batch eviction, time-aware deduplication insertion, and lazy compaction. We implemented and evaluated STsCache via benchmarks and production environments. STsCache can increase throughput of popular time-series databases (InfluxDB, TimescaleDB) by 4.8–10.8X and reduce latency by 79.9%-93.5%. Compared with the latest time-series caching schemes (TSCache, BSCache), STsCache can increase throughput by 1.5–4.5X, reduce latency by 59.4%-81.9%, and increase hit ratios by 22.5%-82.4%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- What Modern NVMe Storage Can Do, And How To Exploit It: High-Performance I/O for High-Performance Storage EnginesGabriel Haas, Viktor LeisVLDB 2023 · 被引用 83 次
- TS-Benchmark: A Benchmark for Time Series DatabasesYuanzhe Hao, Xiongpai Qin, Yueguo Chen, Yaru Li 等ICDE 2021 · 被引用 44 次
- TreeLine: An Update-In-Place Key-Value Store for Modern StorageGeoffrey X. Yu, Markos Markakis, Andreas Kipf, Per-Åke Larson 等VLDB 2023 · 被引用 36 次
- TSM-Bench: Benchmarking Time Series Database Systems for Monitoring ApplicationsAbdelouahab Khelifati, Mourad Khayati, Anton Dignös, Djellel Eddine Difallah 等VLDB 2023 · 被引用 24 次
- Tri-Level Navigator: LLM-Empowered Tri-Level Learning for Time Series OOD GeneralizationChengtao Jian, Kai Yang, Yang JiaoNeurIPS 2024 · 被引用 19 次
相关 Paper
- TSCache: An Efficient Flash-based Caching Scheme for Time-series Data WorkloadsJian Liu, Kefei Wang, Feng ChenVLDB 2021 · 被引用 12 次
- Heracles: An Efficient Storage Model And Data Flushing For Performance Monitoring TimeseriesZhiqi Wang, Jin Xue, Zili ShaoVLDB 2021 · 被引用 14 次
- A large scale analysis of hundreds of in-memory cache clusters at TwitterJuncheng Yang, Yao Yue, K. V. RashmiOSDI 2020 · 被引用 245 次
- Visualization-aware Time Series Min-Max Caching with Error Bound GuaranteesStavros Maroulis, Vassilis Stamatopoulos, George Papastefanatos, Manolis TerrovitisVLDB 2024 · 被引用 8 次
- Competitive Consistent Caching for TransactionsShuai An, Yang Cao, Wenyue ZhaoICDE 2022 · 被引用 1 次
