SEMixer: Semantics Enhanced MLP-Mixer for Multiscale Mixing and Long-term Time Series Forecasting
Xu Zhang, Qitong Wang, Peng Wang, Wei Wang
Abstract
Modeling multiscale patterns is crucial for long-term time series forecasting (TSF). However, redundancy and noise in time series, together with semantic gaps between non-adjacent scales, make the efficient alignment and integration of multi-scale temporal dependencies challenging. To address this, we propose SEMixer, a lightweight multiscale model designed for long-term TSF. SEMixer features two key components: a Random Attention Mechanism (RAM) and a Multiscale Progressive Mixing Chain (MPMC). RAM captures diverse time-patch interactions during training and aggregates them via dropout ensemble at inference, enhancing patchlevel semantics and enabling MLP-Mixer to better model multiscale dependencies. MPMC further stacks RAM and MLP-Mixer in a memory-efficient manner, achieving more effective temporal mixing. It addresses semantic gaps across scales and facilitates better multiscale modeling and forecasting performance. We not only validate the effectiveness of SEMixer on 10 public datasets, but also on the 2025 CCF AlOps Challenge based on 21GB real wireless network data, where SEMixer achieves third place. The code is available at https://github.com/Meteor-Stars/SEMixer . CCS Concepts • Information systems → Spatial-temporal systems; • Computing methodologies → Machine learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9c408d5f-08ad-427f-866f-193345bd4e8eBuilds on27
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang et al.AAAI 2021 · 7,289 citations
- Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series ForecastingHaixu Wu, Jiehui Xu, Jianmin Wang, Mingsheng LongNeurIPS 2021 · 5,824 citations
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- Are Transformers Effective for Time Series Forecasting?Ailing Zeng, Muxi Chen, Lei Zhang, Qiang XuAAAI 2023 · 3,619 citations
- FEDformer: Frequency Enhanced Decomposed Transformer for Long-term Series ForecastingTian Zhou, Ziqing Ma, Qingsong Wen, Xue Wang et al.ICML 2022 · 2,912 citations
Related papers
- TSMixer: Lightweight MLP-Mixer Model for Multivariate Time Series ForecastingVijay Ekambaram, Arindam Jati, Nam Nguyen, Phanwadee Sinthong et al.KDD 2023 · 221 citations
- HDMixer: Hierarchical Dependency with Extendable Patch for Multivariate Time Series ForecastingQihe Huang, Lei Shen, Ruixin Zhang, Jiahuan Cheng et al.AAAI 2024 · 91 citations
- TimeMixer: Decomposable Multiscale Mixing for Time Series ForecastingShiyu Wang, Haixu Wu, Xiaoming Shi, Tengge Hu et al.ICLR 2024 · 573 citations
- Adaptive Multi-Scale Decomposition Framework for Time Series ForecastingYifan Hu, Peiyuan Liu, Peng Zhu, Dawei Cheng et al.AAAI 2025 · 60 citations
- WPMixer: Efficient Multi-Resolution Mixing for Long-Term Time Series ForecastingMd Mahmuddun Nabi Murad, Mehmet Aktukmak, Yasin YilmazAAAI 2025 · 17 citations
