Dynamic TMoE: A Drift-Aware Dynamic Mixture of Experts Framework for Non-Stationary Time Series Forecasting
Jiawen Zhu, Shuhan Liu, Di Weng, Yingcai Wu
Abstract
Non-stationary time series forecasting is challenged by evolving distribution shifts that static models struggle to capture. While Mixture-of-Experts (MoE) architectures offer a promising paradigm for decoupling complex drift patterns, existing approaches are limited by fixed expert pools and memoryless routing, hampering their ability to adapt to abrupt regime shifts. To address this, we propose Dynamic TMoE , a framework that unifies architectural evolution with temporal continuity during learning phase. By detecting distribution shifts via Maximum Mean Discrepancy (MMD), we dynamically instantiate heterogeneous experts and prune redundant ones to optimize capacity. Additionally, a temporal memory router leverages recurrent states and an anomaly repository to ensure stable, context-aware expert selection without requiring test-time updates. Experiments on nine benchmarks demonstrate state-of-the-art performance, reducing MSE by 10.4% and MAE by 7.8%. Code is available at https://github.com/andone-07/Dynamic-TMoE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3cdac34e-d632-42e0-957f-8d7f6dacc892Builds on18
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang et al.AAAI 2021 · 7,289 citations
- Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series ForecastingHaixu Wu, Jiehui Xu, Jianmin Wang, Mingsheng LongNeurIPS 2021 · 5,824 citations
- Are Transformers Effective for Time Series Forecasting?Ailing Zeng, Muxi Chen, Lei Zhang, Qiang XuAAAI 2023 · 3,619 citations
- FEDformer: Frequency Enhanced Decomposed Transformer for Long-term Series ForecastingTian Zhou, Ziqing Ma, Qingsong Wen, Xue Wang et al.ICML 2022 · 2,912 citations
- Reversible Instance Normalization for Accurate Time-Series Forecasting against Distribution ShiftTaesung Kim, Jinhee Kim, Yunwon Tae, Cheonbok Park et al.ICLR 2022 · 1,020 citations
Related papers
- Dynamic Multi-period Experts for Online Time Series ForecastingSeungha Hong, Sukang Chae, Suyeon Kim, Sanghwan Jang et al.WWW 2026
- Mixture of Online and Offline Experts for Non-Stationary Time SeriesZhilin Zhao, Longbing Cao, Yuan-Yu WanAAAI 2025 · 1 citation
- Meta-DMoE: Adapting to Domain Shift by Meta-Distillation from Mixture-of-ExpertsTao Zhong, Zhixiang Chi, Li Gu, Yang Wang et al.NeurIPS 2022 · 70 citations
- Mixture of Prototypes for Test-time Adaptive SegmentationGuangrui Li, Zhengyu Zhu, Yongxin GeCVPR 2026 · 1 citation
- Test-Time Mixture of World Models for Embodied Agents in Dynamic EnvironmentsJinwoo Jang, Minjong Yoo, Sihyung Yoon, Honguk WooICLR 2026 · 2 citations
