S^2-KD: Semantic-Spectral Knowledge Distillation Spatiotemporal Forecasting
Wenshuo Wang, Yaomin Shen, Yingjie Tan, Yihao Chen
Abstract
Spatiotemporal forecasting often relies on computationally intensive models to capture complex dynamics. Knowledge distillation (KD) has emerged as a key technique for creating lightweight student models, with recent advances like frequency-aware KD successfully preserving spectral properties (i.e., high-frequency details and low-frequency trends). However, these methods are fundamentally constrained by operating on pixel-level signals, leaving them blind to the rich semantic and causal context behind the visual patterns. To overcome this limitation, we introduce S 2 -KD, a novel framework that unifies Semantic priors with Spectral representations for distillation. Our approach begins by training a privileged, multimodal teacher model. This teacher leverages textual narratives from a Large Multimodal Model (LMM) to reason about the underlying causes of events, while its architecture simultaneously decouples spectral components in its latent space. The core of our framework is a new distillation objective that transfers this unified semantic-spectral knowledge into a lightweight, vision-only student. Consequently, the student learns to make predictions that are not only spectrally accurate but also semantically coherent, without requiring any textual input or architectural overhead at inference. Extensive experiments on benchmarks like Weather-Bench and TaxiBJ+ show that S 2 -KD significantly boosts the performance of simple student models, enabling them to outperform state-of-the-art methods, particularly in long-horizon and complex non-stationary scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 24629697-ae36-4162-a0c3-2c9c5e7023d4Builds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Fourier Neural Operator for Parametric Partial Differential EquationsZongyi Li, Nikola Borislavov Kovachki, Kamyar Azizzadenesheli, Burigede Liu et al.ICLR 2021 · 3,911 citations
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine et al.AAAI 2020 · 1,361 citations
Related papers
- Frequency-Aligned Knowledge Distillation for Lightweight Spatiotemporal ForecastingYuqi Li, Chuanguang Yang, Hansheng Zeng, Zeyu Dong et al.ICCV 2025 · 23 citations
- Multi-modal Knowledge Distillation-based Human Trajectory ForecastingJaewoo Jeong, Seohee Lee, Daehee Park, Giwon Lee et al.CVPR 2025
- Entropy-Monitored Kernelized Token Distillation for Audio-Visual CompressionHyoungseob Park, Lipeng Ke, Pritish Mohapatra, Huajun Ying et al.ICLR 2026
- OccamVTS: Distilling Vision Models to 1% Parameters for Time Series ForecastingSisuo Lyu, Siru Zhong, Weilin Ruan, Qingxiang Liu et al.AAAI 2026
- Generative Model-Based Feature Knowledge Distillation for Action RecognitionGuiqin Wang, Peng Zhao, Yanjiang Shi, Cong Zhao et al.AAAI 2024 · 9 citations
