Online Multi-Modal Spatio-Temporal Prediction: a Reinforcement Learning and Dynamic Contrastive Framework
Ziquan Fang, Tinghui Luo, Xiaole Pan, Lu Chen, Surun Ji, Mingfan Lu
Abstract
Spatio-temporal prediction is fundamental for traffic management, environmental monitoring, and weather forecasting. While integrating multi-modal information can substantially improve predictive accuracy, the dynamic and evolving nature of spatio-temporal data poses three major challenges: (i) the necessity to adaptively adjust modality contributions as data distributions shift over time; (ii) the presence of sensor noise and unreliable modalities, which undermine prediction robustness; and (iii) the computational overhead of large multi-modal architectures that limits online deployment. To overcome these issues, we propose ROMST, a Reinforcement Learning and Dynamic Contrastive Framework for Online Multi-Modal Spatio-Temporal Prediction. ROMST introduces a reinforcement learning-based adaptive fusion mechanism that continuously optimizes inter-modal weights under streaming and non-stationary conditions. Besides, a Dynamic Contrastive Learning (DCL) module exploits spatio-temporal correlations to distinguish informative signals from noise, improving robustness in dynamic environments. To enhance efficiency, ROMST leverages the Mamba state-space architecture for linear-complexity sequence modeling and applies unstructured pruning to large language models (LLMs) for lightweight textual encoding. Extensive experiments on four real-world multi-modal datasets demonstrate that ROMST consistently outperforms state-of-the-art baselines, achieving up to 37.06% improvement in accuracy and 72.15% reduction in computational cost. The source code and datasets are publicly available at https://github.com/ZJU-DAILY/ROMST.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get a984078d-9eee-4ab3-bc7e-cb56edd612bbRelated papers
- Towards Online Spatio-Temporal Prediction: A Knowledge Distillation Driven Continual Learning ApproachTinghui Luo, Ziquan Fang, Kaixuan Duan, Lu Chen et al.ICDE 2025 · 4 citations
- ST-VLM: A Spatial-to-Image Multimodal Spatial-Temporal Prediction Framework with Vision-Language ModelTong Zhao, Junping Du, Zhe Xue, Meiyu Liang et al.AAAI 2026
- STM3: Mixture of Multiscale Mamba for Long-Term Spatio-Temporal Time-Series PredictionHaolong Chen, Liang Zhang, Zhengyuan Xin, Guangxu ZhuKDD 2026 · 1 citation
- Causal Spatio-Temporal Prediction: An Effective and Efficient Multi-Modal ApproachYuting Huang, Ziquan Fang, Zhihao Zeng, Lu Chen et al.NeurIPS 2025 · 6 citations
- Adaptive Multimodal Fusion: Dynamic Attention Allocation for Intent RecognitionBo Hu, Kai Zhang, Yanghai Zhang, Yuyang YeAAAI 2025 · 6 citations
