Segment, Shuffle, and Stitch: A Simple Layer for Improving Time-Series Representations
Shivam Grover, Amin Jalali, Ali Etemad
Abstract
Existing approaches for learning representations of time-series keep the temporal arrangement of the time-steps intact with the presumption that the original order is the most optimal for learning. However, non-adjacent sections of real-world time-series may have strong dependencies. Accordingly, we raise the question: Is there an alternative arrangement for time-series which could enable more effective representation learning? To address this, we propose a simple plug-and-play neural network layer called Segment, Shuffle, and Stitch (S3) designed to improve representation learning in time-series models. S3 works by creating non-overlapping segments from the original sequence and shuffling them in a learned manner that is optimal for the task at hand. It then re-attaches the shuffled segments back together and performs a learned weighted sum with the original input to capture both the newly shuffled sequence along with the original sequence. S3 is modular and can be stacked to achieve different levels of granularity, and can be added to many forms of neural architectures including CNNs or Transformers with negligible computation overhead. Through extensive experiments on several datasets and state-of-the-art baselines, we show that incorporating S3 results in significant improvements for the tasks of time-series classification, forecasting, and anomaly detection, improving performance on certain datasets by up to 68%. We also show that S3 makes the learning more stable with a smoother training loss curve and loss landscape compared to the original baseline. The code is available at https://github.com/shivam-grover/S3-TimeSeries.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4f38082e-4800-4552-8747-e51c0f4c9badCited by top-tier papers1
Ask how each one uses itBuilds on27
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang et al.AAAI 2021 · 7,289 citations
- Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series ForecastingHaixu Wu, Jiehui Xu, Jianmin Wang, Mingsheng LongNeurIPS 2021 · 5,824 citations
- Are Transformers Effective for Time Series Forecasting?Ailing Zeng, Muxi Chen, Lei Zhang, Qiang XuAAAI 2023 · 3,619 citations
- Adaptive Graph Convolutional Recurrent Network for Traffic ForecastingLei Bai, Lina Yao, Can Li, Xianzhi Wang et al.NeurIPS 2020 · 2,206 citations
- TS2Vec: Towards Universal Representation of Time SeriesZhihan Yue, Yujing Wang, Juanyong Duan, Tianmeng Yang et al.AAAI 2022 · 938 citations
Related papers
- Learning to Embed Time Series Patches IndependentlySeunghan Lee, Taeyoung Park, Kibok LeeICLR 2024 · 57 citations
- T-Rep: Representation Learning for Time Series using Time-EmbeddingsArchibald Fraikin, Adrien Bennetot, Stéphanie AllassonnièreICLR 2024 · 24 citations
- TSLANet: Rethinking Transformers for Time Series Representation LearningEmadeldeen Eldele, Mohamed Ragab, Zhenghua Chen, Min Wu et al.ICML 2024 · 159 citations
- Enhancing Time Series Forecasting through Selective Representation Spaces: A Patch PerspectiveXingjian Wu, Xiangfei Qiu, Hanyin Cheng, Zhengyu Li et al.NeurIPS 2025 · 58 citations
- PaAno: Patch-Based Representation Learning for Time-Series Anomaly DetectionJinju Park, Seokho KangICLR 2026 · 6 citations
