Does Vector Quantization Fail in Spatio-Temporal Forecasting? Exploring a Differentiable Sparse Soft-Vector Quantization Approach
Chao Chen, Tian Zhou, Yanjun Zhao, Hui Liu, Rong Jin, Liang Sun
摘要
Spatio-temporal forecasting is crucial in various fields and requires a careful balance between identifying subtle patterns and filtering out noise. Vector quantization (VQ) appears well-suited for this purpose, as it quantizes input vectors into a set of codebook vectors or patterns. Although VQ has shown promise in various computer vision tasks, it surprisingly falls short in enhancing the accuracy of spatio-temporal forecasting. We attribute this to two main issues: inaccurate optimization due to non-differentiability and limited representation power in hard-VQ. To tackle these challenges, we introduce Differentiable Sparse Soft-Vector Quantization (SVQ), the first VQ method to enhance spatio-temporal forecasting. SVQ balances detail preservation with noise reduction, offering full differentiability and a solid foundation in sparse regression. Our approach employs a two-layer MLP and an extensive codebook to streamline the sparse regression process, significantly cutting computational costs while simplifying training and improving performance. Empirical studies on five spatio-temporal benchmark datasets show SVQ achieves state-of-the-art results, including a 7.9% improvement on the WeatherBench-S temperature dataset and an average mean absolute error reduction of 9.4% in video prediction benchmarks (Human3.6M, KTH, and KittiCaltech), along with a 17.3% enhancement in image quality (LPIPS). Code is publicly available at https://github.com/Pachark/SVQ-Forecasting . CCS Concepts • Computing methodologies → Supervised learning by regression.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- SIGMA-PPG: Statistical-prior Informed Generative Masking Architecture for PPG Foundation ModelZONGHENG GUO, Tao Chen, Yang Jiao, Yi Pan 等ICML 2026 · 被引用 3 次
- PhyOceanCast: Global Ocean Forecasting with Physics-Informed DiffusionQixiu Li, Xiang Zhu, Xiaoyong Li, Xiaolong XuCVPR 2026
它引用的顶会 Paper16
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer 等NeurIPS 2021 · 被引用 3,862 次
- MetaFormer is Actually What You Need for VisionWeihao Yu, Mi Luo, Pan Zhou, Chenyang Si 等CVPR 2022 · 被引用 1,114 次
- Language Model Beats Diffusion - Tokenizer is key to visual generationLijun Yu, José Lezama, Nitesh Bharadwaj Gundavarapu, Luca Versari 等ICLR 2024 · 被引用 609 次
相关 Paper
- Moment Quantization for Video Temporal GroundingXiaolong Sun, Le Wang, Sanping Zhou, Liushuai Shi 等ICCV 2025 · 被引用 2 次
- Adverse Weather Removal with Codebook PriorsTian Ye, Sixiang Chen, Jinbin Bai, Jun Shi 等ICCV 2023 · 被引用 64 次
- PredToken: Predicting Unknown Tokens and Beyond with Coarse-to-Fine Iterative DecodingXuesong Nie, Haoyuan Jin, Yunfeng Yan, Xi Chen 等CVPR 2024
- Decoupled Spatiotemporal Forecasting from Extreme Sparse Observations via Quantized Latent SpaceZhongnan Weng, Yue Hong, Hang Yu, Jiayi Que 等AAAI 2026
- 3D Self-Attention for Unsupervised Video QuantizationJingkuan Song, Ruimin Lang, Xiaosu Zhu, Xing Xu 等SIGIR 2020 · 被引用 3 次
