Does Vector Quantization Fail in Spatio-Temporal Forecasting? Exploring a Differentiable Sparse Soft-Vector Quantization Approach
Chao Chen, Tian Zhou, Yanjun Zhao, Hui Liu, Rong Jin, Liang Sun
Abstract
Spatio-temporal forecasting is crucial in various fields and requires a careful balance between identifying subtle patterns and filtering out noise. Vector quantization (VQ) appears well-suited for this purpose, as it quantizes input vectors into a set of codebook vectors or patterns. Although VQ has shown promise in various computer vision tasks, it surprisingly falls short in enhancing the accuracy of spatio-temporal forecasting. We attribute this to two main issues: inaccurate optimization due to non-differentiability and limited representation power in hard-VQ. To tackle these challenges, we introduce Differentiable Sparse Soft-Vector Quantization (SVQ), the first VQ method to enhance spatio-temporal forecasting. SVQ balances detail preservation with noise reduction, offering full differentiability and a solid foundation in sparse regression. Our approach employs a two-layer MLP and an extensive codebook to streamline the sparse regression process, significantly cutting computational costs while simplifying training and improving performance. Empirical studies on five spatio-temporal benchmark datasets show SVQ achieves state-of-the-art results, including a 7.9% improvement on the WeatherBench-S temperature dataset and an average mean absolute error reduction of 9.4% in video prediction benchmarks (Human3.6M, KTH, and KittiCaltech), along with a 17.3% enhancement in image quality (LPIPS). Code is publicly available at https://github.com/Pachark/SVQ-Forecasting . CCS Concepts • Computing methodologies → Supervised learning by regression.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- SIGMA-PPG: Statistical-prior Informed Generative Masking Architecture for PPG Foundation ModelZONGHENG GUO, Tao Chen, Yang Jiao, Yi Pan et al.ICML 2026 · 3 citations
- PhyOceanCast: Global Ocean Forecasting with Physics-Informed DiffusionQixiu Li, Xiang Zhu, Xiaoyong Li, Xiaolong XuCVPR 2026
Builds on16
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- MetaFormer is Actually What You Need for VisionWeihao Yu, Mi Luo, Pan Zhou, Chenyang Si et al.CVPR 2022 · 1,114 citations
- Language Model Beats Diffusion - Tokenizer is key to visual generationLijun Yu, José Lezama, Nitesh Bharadwaj Gundavarapu, Luca Versari et al.ICLR 2024 · 609 citations
Related papers
- Moment Quantization for Video Temporal GroundingXiaolong Sun, Le Wang, Sanping Zhou, Liushuai Shi et al.ICCV 2025 · 2 citations
- Adverse Weather Removal with Codebook PriorsTian Ye, Sixiang Chen, Jinbin Bai, Jun Shi et al.ICCV 2023 · 64 citations
- PredToken: Predicting Unknown Tokens and Beyond with Coarse-to-Fine Iterative DecodingXuesong Nie, Haoyuan Jin, Yunfeng Yan, Xi Chen et al.CVPR 2024
- Decoupled Spatiotemporal Forecasting from Extreme Sparse Observations via Quantized Latent SpaceZhongnan Weng, Yue Hong, Hang Yu, Jiayi Que et al.AAAI 2026
- 3D Self-Attention for Unsupervised Video QuantizationJingkuan Song, Ruimin Lang, Xiaosu Zhu, Xing Xu et al.SIGIR 2020 · 3 citations
