TimeDART: A Diffusion Autoregressive Transformer for Self-Supervised Time Series Representation
Daoyu Wang, Mingyue Cheng, Zhiding Liu, Qi Liu
Abstract
Self-supervised learning has garnered increasing attention in time series analysis for benefiting various downstream tasks and reducing reliance on labeled data. Despite its effectiveness, existing methods often struggle to comprehensively capture both long-term dynamic evolution and subtle local patterns in a unified manner. In this work, we propose TimeDART, a novel self-supervised time series pre-training framework that unifies two powerful generative paradigms to learn more transferable representations. Specifically, we first employ a causal Transformer encoder, accompanied by a patch-based embedding strategy, to model the evolving trends from left to right. Building on this global modeling, we further introduce a denoising diffusion process to capture fine-grained local patterns through forward diffusion and reverse denoising. Finally, we optimize the model in an autoregressive manner. As a result, TimeDART effectively accounts for both global and local sequence features in a coherent way. We conduct extensive experiments on public datasets for time series forecasting and classification. The experimental results demonstrate that TimeDART consistently outperforms previous compared methods, validating the effectiveness of our approach. Our code is available at https://github.com/Melmaphother/TimeDART .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- MemCast: Memory-Driven Time Series Forecasting with Experience-Conditioned ReasoningXiaoyu Tao, Mingyue Cheng, Ze Guo, Shuo Yu et al.ICML 2026 · 8 citations
- CoGenCast: A Coupled Autoregressive–Flow Generative Framework for Time Series ForecastingMingyue Cheng, Yaguo Liu, Daoyu Wang, Xiaoyu Tao et al.ICML 2026
Builds on25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang et al.AAAI 2021 · 7,289 citations
- Are Transformers Effective for Time Series Forecasting?Ailing Zeng, Muxi Chen, Lei Zhang, Qiang XuAAAI 2023 · 3,619 citations
- Reversible Instance Normalization for Accurate Time-Series Forecasting against Distribution ShiftTaesung Kim, Jinhee Kim, Yunwon Tae, Cheonbok Park et al.ICLR 2022 · 1,020 citations
Related papers
- RePatch: Learning Entropy-Guided Patch Structures with Quantized Representations for Time Series ForecastingHanbin Xiao, Xun Zhou, Rui Huang, Xiucheng Li et al.KDD 2026
- TimeCAP: A Channel-Aware Pre-Training Framework for Multivariate Time Series ForecastingChuanru Ren, Yao Lu, Tianjin Huang, Haowen Zheng et al.AAAI 2026
- A Time Series is Worth 64 Words: Long-term Forecasting with TransformersYuqi Nie, Nam H. Nguyen, Phanwadee Sinthong, Jayant KalagnanamICLR 2023 · 536 citations
- Multi-Patch Prediction: Adapting Language Models for Time Series Representation LearningYuxuan Bian, Xuan Ju, Jiangtong Li, Zhijian Xu et al.ICML 2024 · 13 citations
- T-Rep: Representation Learning for Time Series using Time-EmbeddingsArchibald Fraikin, Adrien Bennetot, Stéphanie AllassonnièreICLR 2024 · 24 citations
