Wavelet-Driven Spatiotemporal Predictive Learning: Bridging Frequency and Time Variations
Xuesong Nie, Yunfeng Yan, Siyuan Li, Cheng Tan, Xi Chen, Haoyuan Jin, Zhihang Zhu, Stan Z. Li, Donglian Qi
摘要
Spatiotemporal predictive learning is a paradigm that empowers models to learn spatial and temporal patterns by predicting future frames from past frames in an unsupervised manner. This method typically uses recurrent units to capture long-term dependencies, but these units often come with high computational costs and limited performance in real-world scenes. This paper presents an innovative Wavelet-based Spa-tioTemporal (WaST) framework, which extracts and adaptively controls both low and high-frequency components at image and feature levels via 3D discrete wavelet transform for faster processing while maintaining high-quality predictions. We propose a Time-Frequency Aware Translator uniquely crafted to efficiently learn short-and long-range spatiotemporal information by individually modeling spatial frequency and temporal variations. Meanwhile, we design a wavelet-domain High-Frequency Focal Loss that effectively supervises high-frequency variations. Extensive experiments across various real-world scenarios, such as driving scene prediction, traffic flow prediction, human motion capture, and weather forecasting, demonstrate that our proposed WaST achieves state-of-the-art performance over various spatiotemporal prediction methods. Our code is available at https://github.com/xuesongnie/WaST.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- PFGNet: A Fully Convolutional Frequency-Guided Peripheral Gating Network for Efficient Spatiotemporal Predictive LearningXinyong Cai, Changbin Sun, Yong Wang, Hongyu Yang 等CVPR 2026 · 被引用 3 次
- Met2Net: A Decoupled Two-Stage Spatio-Temporal Forecasting Model for Complex Meteorological SystemsShaohan Li, Hao Yang, Min Chen, Xiaolin QinICCV 2025 · 被引用 1 次
- UniSTD: Towards Unified Spatio-Temporal Learning across Diverse DisciplinesChen Tang, Xinzhu Ma, Encheng Su, Xiufeng Song 等CVPR 2025
它引用的顶会 Paper9
- MetaFormer is Actually What You Need for VisionWeihao Yu, Mi Luo, Pan Zhou, Chenyang Si 等CVPR 2022 · 被引用 1,114 次
- Focal Frequency Loss for Image Reconstruction and SynthesisLiming Jiang, Bo Dai, Wayne Wu, Chen Change LoyICCV 2021 · 被引用 422 次
- SimVP: Simpler yet Better Video PredictionZhangyang Gao, Cheng Tan, Lirong Wu, Stan Z. LiCVPR 2022 · 被引用 313 次
- MAU: A Motion-Aware Unit for Video Prediction and BeyondZheng Chang, Xinfeng Zhang, Shanshe Wang, Siwei Ma 等NeurIPS 2021 · 被引用 193 次
- Architecture-Agnostic Masked Image Modeling - From ViT back to CNNSiyuan Li, Di Wu, Fang Wu, Zelin Zang 等ICML 2023 · 被引用 60 次
相关 Paper
- WaveForM: Graph Enhanced Wavelet Learning for Long Sequence Forecasting of Multivariate Time SeriesFuhao Yang, Xin Li, Min Wang, Hongyu Zang 等AAAI 2023 · 被引用 35 次
- Temporal Attention Unit: Towards Efficient Spatiotemporal Predictive LearningCheng Tan, Zhangyang Gao, Lirong Wu, Yongjie Xu 等CVPR 2023
- TF-FACE: Time-Frequency Fusion Learning via Frequency-Domain Adaptive and Controllable Enhancement for Trajectory PredictionDongjian Song, Yunhao Meng, Songjun Huang, Jiayi HanICML 2026
- WaveAR: Wavelet-Aware Continuous Autoregressive Diffusion for Accurate Human Motion PredictionShengchuan Gao, Shuo Wang, Yabiao Wang, Ran YiNeurIPS 2025 · 被引用 1 次
- When Spatio-Temporal Meet Wavelets: Disentangled Traffic Forecasting via Efficient Spectral Graph Attention NetworksYuchen Fang, Yanjun Qin, Haiyong Luo, Fang Zhao 等ICDE 2023 · 被引用 134 次
