Scalable Time Series Compound Infrastructure
Noura S. Alghamdi, Liang Zhang, Elke A. Rundensteiner, Mohamed Y. Eltabakh
摘要
Objects ranging from a patient's history of medical tests to an IoT device's series of sensor maintenance records leave digital traces in the form of big time series. These time series objects do not only span exceedingly long time periods (sometimes years), but are also characterized by intermittent yet interrelated time series measurements punctuated by long gaps of silence. This prevalent data type, which we refer to as Time Series Compound objects (or, TSC), has been largely overlooked in the literature. Unique challenges arise when managing, querying and analyzing repositories of these big TSC objects. These include appropriate similarity semantics with time misalignment resiliency, efficient storage of excessively long and complex objects, and TSC-holistic indexing. We demonstrate that state-of-the-art time series systems, although effective at indexing and searching regular time series data, fail to support such big TSC data. In this work, we introduce the first comprehensive solution for managing TSC objects as first class citizen. We introduce new similarity-match semantics as well as a compact misalignment-resilient representation for TSCs. Upon this foundation, we then design a TSC-aware distributed indexing infrastructure Sloth that supports scalable storage, indexing and querying of TB-scale TSC datasets. Our experimental study demonstrates that for TB-scale datasets, the query response time of Sloth is up to one order of magnitude faster than that of existing systems, while the mean average precision (mAP) for approximate kNN similarity match query results by Sloth is 70% more accurate than existing solutions.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- CLIMBER: Pivot-Based Approximate Similarity Search Over Big Data SeriesLiang Zhang, Mohamed Y. Eltabakh, Elke A. Rundensteiner, Khalid AlnuaimICDE 2024 · 被引用 1 次
- In-Database Time Series ClusteringYunxiang Su, Kenny Ye Liang, Shaoxu SongSIGMOD 2025 · 被引用 3 次
- ChainLink: Indexing Big Time Series Data For Long Subsequence MatchingNoura Alghamdi, Liang Zhang, Huayi Zhang, Elke A. Rundensteiner 等ICDE 2020 · 被引用 15 次
- STsCache: An Efficient Semantic Caching Scheme for Time-series Data Workloads Based on Hybrid StorageTao Kong, Hui Li, Yuxuan Zhao, Liping Li 等VLDB 2025
- ForestTI: A Scalable Inverted-Index-Oriented Timeseries Management System with Flexible Memory EfficiencyZhiqi Wang, Zili ShaoSIGMOD 2023 · 被引用 2 次
