SPARTAN: Data-Adaptive Symbolic Time-Series Approximation
Fan Yang, John Paparrizos
Abstract
Symbolic approximations are dimensionality reduction techniques that convert time series into sequences of discrete symbols, enhancing interpretability while reducing computational and storage costs. To construct symbolic representations, first numeric representations approximate and capture properties of raw time series, followed by a discretization step that converts these numeric dimensions into symbols. Despite decades of development, existing approaches have several key limitations that often result in unsatisfactory performance: they (i) rely on data-agnostic numeric approximations, disregarding intrinsic properties of the time series; (ii) decompose dimensions into equal-sized subspaces, assuming independence among dimensions; and (iii) allocate a uniform encoding budget for discretizing each dimension or subspace, assuming balanced importance. To address these shortcomings, we propose SPARTAN, a novel data-adaptive symbolic approximation method that intelligently allocates the encoding budget according to the importance of the constructed uncorrelated dimensions. Specifically, SPARTAN (i) leverages intrinsic dimensionality reduction properties to derive non-overlapping, uncorrelated latent dimensions; (ii) adaptively distributes the budget based on the importance of each dimension by solving a constrained optimization problem; and (iii) prevents false dismissals in similarity search by ensuring a lower bound on the true distance in the original space. To demonstrate SPARTAN's robustness, we conduct the most comprehensive study to date, comparing SPARTAN with seven state-of-the-art symbolic methods across four tasks: classification, clustering, indexing, and anomaly detection. Rigorous statistical analysis across hundreds of datasets shows that SPARTAN outperforms competing methods significantly on all tasks in terms of downstream accuracy, given the same budget. Notably, SPARTAN achieves up to a 2x speedup compared to the most accurate rival. Overall, SPARTAN effectively improves the symbolic representation quality without storage or runtime overheads, paving the way for future advancements.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0a34ce52-67de-4c92-94e5-74ada5392a97Cited by top-tier papers9
- Understanding the Black Box: A Deep Empirical Dive into Shapley Value Approximations for Tabular DataSuchit Gupte, John PaparrizosSIGMOD 2025 · 19 citations
- A Structured Study of Multivariate Time-Series Distance MeasuresJens E. d'Hondt, Haojun Li, Fan Yang, Odysseas Papapetrou et al.SIGMOD 2025 · 14 citations
- TSB-AutoAD: Towards Automated Solutions for Time-Series Anomaly Detection [E, A & B]Qinghua Liu, Seunghak Lee, John PaparrizosVLDB 2025 · 13 citations
- Time-Series Clustering: A Comprehensive Study of Data Mining, Machine Learning, and Deep Learning MethodsJohn Paparrizos, Bogireddy Sai Prasanna TejaVLDB 2025 · 13 citations
- Beyond Compression: A Comprehensive Evaluation of Lossless Floating-Point CompressionKaisei Hishida, Chunwei Liu, John Paparrizos, Aaron J. ElmoreVLDB 2025 · 8 citations
Builds on26
- One Fits All: Power General Time Series Analysis by Pretrained LMTian Zhou, Peisong Niu, Xue Wang, Liang Sun et al.NeurIPS 2023 · 1,178 citations
- TS2Vec: Towards Universal Representation of Time SeriesZhihan Yue, Yujing Wang, Juanyong Duan, Tianmeng Yang et al.AAAI 2022 · 938 citations
- Time-LLM: Time Series Forecasting by Reprogramming Large Language ModelsMing Jin, Shiyu Wang, Lintao Ma, Zhixuan Chu et al.ICLR 2024 · 915 citations
- Large Language Models Are Zero-Shot Time Series ForecastersNate Gruver, Marc Finzi, Shikai Qiu, Andrew Gordon WilsonNeurIPS 2023 · 898 citations
- Self-Supervised Contrastive Pre-Training For Time Series via Time-Frequency ConsistencyXiang Zhang, Ziyuan Zhao, Theodoros Tsiligkaridis, Marinka ZitnikNeurIPS 2022 · 558 citations
Related papers
- Fast and Exact Similarity Search in Less than a Blink of an EyePatrick Schäfer, Jakob Brand, Ulf Leser, Botao Peng et al.ICDE 2025 · 1 citation
- Fast Adaptive Similarity Search through Variance-Aware QuantizationJohn Paparrizos, Ikraduya Edian, Chunwei Liu, Aaron J. Elmore et al.ICDE 2022 · 34 citations
- Synthetic Series-Symbol Data Generation for Time Series Foundation ModelsWenxuan Wang, Kai Wu, Yujian Betterest Li, Dan Wang et al.NeurIPS 2025 · 1 citation
- Sim-Piece: Highly Accurate Piecewise Linear Approximation through Similar Segment MergingXenophon Kitsios, Panagiotis Liakos, Katia Papakonstantinopoulou, Yannis KotidisVLDB 2023 · 20 citations
- leSAX Index: A Learned SAX Representation Index for Time Series Similarity SearchGuozhong Li, Byron Choi, Rundong Zuo, Sourav S. Bhowmick et al.ICDE 2025 · 3 citations
