ST-VLM: A Spatial-to-Image Multimodal Spatial-Temporal Prediction Framework with Vision-Language Model
Tong Zhao, Junping Du, Zhe Xue, Meiyu Liang, Aijing Li, Xiaolong Meng, Dandan Liu
Abstract
Spatial-temporal prediction plays a crucial role in various domains, including intelligent transportation and environmental monitoring. Although large language model has shown advantages in long-range dependency modeling and excellent generalization ability for prediction tasks, it has limited understanding of spatial-temporal features. Especially for spatial features, most existing methods still simplify the spatialtemporal prediction task into multiple independent temporal prediction tasks, failing to effectively encode the dynamic evolution of spatial relations. To address these problems, we propose ST-VLM (Spatial-Temporal Forecasting with Vision-Language Model), a novel framework that leverages visual representations to encode the dynamic spatial dependencies within spatial-temporal data and integrates multi-modal information to enhance prediction. This framework transforms spatial-temporal features into three modalities: vision, text, and time series, enhances cross-modal fusion through an attention-aware fusion mechanism in the first-layer of Vision-Language Model (VLM), optimizes multi-modal feature interaction via adaptive fine-tuning strategies. After fusion, the multi-modal embeddings are subsequently used for the final spatial-temporal prediction task. Extensive experiments demonstrate that ST-VLM achieves state-of-the-art performance across various datasets. In particular, the framework exhibits promising results in few-shot scenarios, verifying its strong generalization ability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext af81e404-1c17-4c01-93fc-adf6a36554c7Builds on16
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
- Adaptive Graph Convolutional Recurrent Network for Traffic ForecastingLei Bai, Lina Yao, Can Li, Xianzhi Wang et al.NeurIPS 2020 · 2,206 citations
- iTransformer: Inverted Transformers Are Effective for Time Series ForecastingYong Liu, Tengge Hu, Haoran Zhang, Haixu Wu et al.ICLR 2024 · 1,703 citations
- Spatial-Temporal Synchronous Graph Convolutional Networks: A New Framework for Spatial-Temporal Network Data ForecastingChao Song, Youfang Lin, Shengnan Guo, Huaiyu WanAAAI 2020 · 1,659 citations
- One Fits All: Power General Time Series Analysis by Pretrained LMTian Zhou, Peisong Niu, Xue Wang, Liang Sun et al.NeurIPS 2023 · 1,178 citations
Related papers
- Time-VLM: Exploring Multimodal Vision-Language Models for Augmented Time Series ForecastingSiru Zhong, Weilin Ruan, Ming Jin, Huan Li et al.ICML 2025
- FSTLLM: Spatio-Temporal LLM for Few Shot Time Series ForecastingYue Jiang, Yile Chen, Xiucheng Li, Qin Chao et al.ICML 2025
- STEM-LTS: Integrating Semantic-Temporal Dynamics in LLM-driven Time Series AnalysisZhe Zhao, Pengkun Wang, Haibin Wen, Shuang Wang et al.AAAI 2025 · 7 citations
- SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding CapabilityJiankang Wang, Zhihan Zhang, Zhihang Liu, Yang Li et al.AAAI 2026 · 20 citations
- AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document UnderstandingAhmed Masry, Juan A. Rodríguez, Tianyu Zhang, Suyuchen Wang et al.NeurIPS 2025 · 7 citations
