Online Multi-Modal Spatio-Temporal Prediction: a Reinforcement Learning and Dynamic Contrastive Framework
Ziquan Fang, Tinghui Luo, Xiaole Pan, Lu Chen, Surun Ji, Mingfan Lu
摘要
Spatio-temporal prediction is fundamental for traffic management, environmental monitoring, and weather forecasting. While integrating multi-modal information can substantially improve predictive accuracy, the dynamic and evolving nature of spatio-temporal data poses three major challenges: (i) the necessity to adaptively adjust modality contributions as data distributions shift over time; (ii) the presence of sensor noise and unreliable modalities, which undermine prediction robustness; and (iii) the computational overhead of large multi-modal architectures that limits online deployment. To overcome these issues, we propose ROMST, a Reinforcement Learning and Dynamic Contrastive Framework for Online Multi-Modal Spatio-Temporal Prediction. ROMST introduces a reinforcement learning-based adaptive fusion mechanism that continuously optimizes inter-modal weights under streaming and non-stationary conditions. Besides, a Dynamic Contrastive Learning (DCL) module exploits spatio-temporal correlations to distinguish informative signals from noise, improving robustness in dynamic environments. To enhance efficiency, ROMST leverages the Mamba state-space architecture for linear-complexity sequence modeling and applies unstructured pruning to large language models (LLMs) for lightweight textual encoding. Extensive experiments on four real-world multi-modal datasets demonstrate that ROMST consistently outperforms state-of-the-art baselines, achieving up to 37.06% improvement in accuracy and 72.15% reduction in computational cost. The source code and datasets are publicly available at https://github.com/ZJU-DAILY/ROMST.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Towards Online Spatio-Temporal Prediction: A Knowledge Distillation Driven Continual Learning ApproachTinghui Luo, Ziquan Fang, Kaixuan Duan, Lu Chen 等ICDE 2025 · 被引用 4 次
- ST-VLM: A Spatial-to-Image Multimodal Spatial-Temporal Prediction Framework with Vision-Language ModelTong Zhao, Junping Du, Zhe Xue, Meiyu Liang 等AAAI 2026
- STM3: Mixture of Multiscale Mamba for Long-Term Spatio-Temporal Time-Series PredictionHaolong Chen, Liang Zhang, Zhengyuan Xin, Guangxu ZhuKDD 2026 · 被引用 1 次
- Causal Spatio-Temporal Prediction: An Effective and Efficient Multi-Modal ApproachYuting Huang, Ziquan Fang, Zhihao Zeng, Lu Chen 等NeurIPS 2025 · 被引用 6 次
- Adaptive Multimodal Fusion: Dynamic Attention Allocation for Intent RecognitionBo Hu, Kai Zhang, Yanghai Zhang, Yuyang YeAAAI 2025 · 被引用 6 次
