SPARTAN: Data-Adaptive Symbolic Time-Series Approximation
Fan Yang, John Paparrizos
摘要
Symbolic approximations are dimensionality reduction techniques that convert time series into sequences of discrete symbols, enhancing interpretability while reducing computational and storage costs. To construct symbolic representations, first numeric representations approximate and capture properties of raw time series, followed by a discretization step that converts these numeric dimensions into symbols. Despite decades of development, existing approaches have several key limitations that often result in unsatisfactory performance: they (i) rely on data-agnostic numeric approximations, disregarding intrinsic properties of the time series; (ii) decompose dimensions into equal-sized subspaces, assuming independence among dimensions; and (iii) allocate a uniform encoding budget for discretizing each dimension or subspace, assuming balanced importance. To address these shortcomings, we propose SPARTAN, a novel data-adaptive symbolic approximation method that intelligently allocates the encoding budget according to the importance of the constructed uncorrelated dimensions. Specifically, SPARTAN (i) leverages intrinsic dimensionality reduction properties to derive non-overlapping, uncorrelated latent dimensions; (ii) adaptively distributes the budget based on the importance of each dimension by solving a constrained optimization problem; and (iii) prevents false dismissals in similarity search by ensuring a lower bound on the true distance in the original space. To demonstrate SPARTAN's robustness, we conduct the most comprehensive study to date, comparing SPARTAN with seven state-of-the-art symbolic methods across four tasks: classification, clustering, indexing, and anomaly detection. Rigorous statistical analysis across hundreds of datasets shows that SPARTAN outperforms competing methods significantly on all tasks in terms of downstream accuracy, given the same budget. Notably, SPARTAN achieves up to a 2x speedup compared to the most accurate rival. Overall, SPARTAN effectively improves the symbolic representation quality without storage or runtime overheads, paving the way for future advancements.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Understanding the Black Box: A Deep Empirical Dive into Shapley Value Approximations for Tabular DataSuchit Gupte, John PaparrizosSIGMOD 2025 · 被引用 19 次
- A Structured Study of Multivariate Time-Series Distance MeasuresJens E. d'Hondt, Haojun Li, Fan Yang, Odysseas Papapetrou 等SIGMOD 2025 · 被引用 14 次
- TSB-AutoAD: Towards Automated Solutions for Time-Series Anomaly Detection [E, A & B]Qinghua Liu, Seunghak Lee, John PaparrizosVLDB 2025 · 被引用 13 次
- Time-Series Clustering: A Comprehensive Study of Data Mining, Machine Learning, and Deep Learning MethodsJohn Paparrizos, Bogireddy Sai Prasanna TejaVLDB 2025 · 被引用 13 次
- Beyond Compression: A Comprehensive Evaluation of Lossless Floating-Point CompressionKaisei Hishida, Chunwei Liu, John Paparrizos, Aaron J. ElmoreVLDB 2025 · 被引用 8 次
它引用的顶会 Paper26
- One Fits All: Power General Time Series Analysis by Pretrained LMTian Zhou, Peisong Niu, Xue Wang, Liang Sun 等NeurIPS 2023 · 被引用 1,178 次
- TS2Vec: Towards Universal Representation of Time SeriesZhihan Yue, Yujing Wang, Juanyong Duan, Tianmeng Yang 等AAAI 2022 · 被引用 938 次
- Time-LLM: Time Series Forecasting by Reprogramming Large Language ModelsMing Jin, Shiyu Wang, Lintao Ma, Zhixuan Chu 等ICLR 2024 · 被引用 915 次
- Large Language Models Are Zero-Shot Time Series ForecastersNate Gruver, Marc Finzi, Shikai Qiu, Andrew Gordon WilsonNeurIPS 2023 · 被引用 898 次
- Self-Supervised Contrastive Pre-Training For Time Series via Time-Frequency ConsistencyXiang Zhang, Ziyuan Zhao, Theodoros Tsiligkaridis, Marinka ZitnikNeurIPS 2022 · 被引用 558 次
相关 Paper
- Fast and Exact Similarity Search in Less than a Blink of an EyePatrick Schäfer, Jakob Brand, Ulf Leser, Botao Peng 等ICDE 2025 · 被引用 1 次
- Fast Adaptive Similarity Search through Variance-Aware QuantizationJohn Paparrizos, Ikraduya Edian, Chunwei Liu, Aaron J. Elmore 等ICDE 2022 · 被引用 34 次
- Synthetic Series-Symbol Data Generation for Time Series Foundation ModelsWenxuan Wang, Kai Wu, Yujian Betterest Li, Dan Wang 等NeurIPS 2025 · 被引用 1 次
- Sim-Piece: Highly Accurate Piecewise Linear Approximation through Similar Segment MergingXenophon Kitsios, Panagiotis Liakos, Katia Papakonstantinopoulou, Yannis KotidisVLDB 2023 · 被引用 20 次
- leSAX Index: A Learned SAX Representation Index for Time Series Similarity SearchGuozhong Li, Byron Choi, Rundong Zuo, Sourav S. Bhowmick 等ICDE 2025 · 被引用 3 次
