ST-VLM: A Spatial-to-Image Multimodal Spatial-Temporal Prediction Framework with Vision-Language Model
Tong Zhao, Junping Du, Zhe Xue, Meiyu Liang, Aijing Li, Xiaolong Meng, Dandan Liu
摘要
Spatial-temporal prediction plays a crucial role in various domains, including intelligent transportation and environmental monitoring. Although large language model has shown advantages in long-range dependency modeling and excellent generalization ability for prediction tasks, it has limited understanding of spatial-temporal features. Especially for spatial features, most existing methods still simplify the spatialtemporal prediction task into multiple independent temporal prediction tasks, failing to effectively encode the dynamic evolution of spatial relations. To address these problems, we propose ST-VLM (Spatial-Temporal Forecasting with Vision-Language Model), a novel framework that leverages visual representations to encode the dynamic spatial dependencies within spatial-temporal data and integrates multi-modal information to enhance prediction. This framework transforms spatial-temporal features into three modalities: vision, text, and time series, enhances cross-modal fusion through an attention-aware fusion mechanism in the first-layer of Vision-Language Model (VLM), optimizes multi-modal feature interaction via adaptive fine-tuning strategies. After fusion, the multi-modal embeddings are subsequently used for the final spatial-temporal prediction task. Extensive experiments demonstrate that ST-VLM achieves state-of-the-art performance across various datasets. In particular, the framework exhibits promising results in few-shot scenarios, verifying its strong generalization ability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 被引用 2,258 次
- Adaptive Graph Convolutional Recurrent Network for Traffic ForecastingLei Bai, Lina Yao, Can Li, Xianzhi Wang 等NeurIPS 2020 · 被引用 2,206 次
- iTransformer: Inverted Transformers Are Effective for Time Series ForecastingYong Liu, Tengge Hu, Haoran Zhang, Haixu Wu 等ICLR 2024 · 被引用 1,703 次
- Spatial-Temporal Synchronous Graph Convolutional Networks: A New Framework for Spatial-Temporal Network Data ForecastingChao Song, Youfang Lin, Shengnan Guo, Huaiyu WanAAAI 2020 · 被引用 1,659 次
- One Fits All: Power General Time Series Analysis by Pretrained LMTian Zhou, Peisong Niu, Xue Wang, Liang Sun 等NeurIPS 2023 · 被引用 1,178 次
相关 Paper
- Time-VLM: Exploring Multimodal Vision-Language Models for Augmented Time Series ForecastingSiru Zhong, Weilin Ruan, Ming Jin, Huan Li 等ICML 2025
- FSTLLM: Spatio-Temporal LLM for Few Shot Time Series ForecastingYue Jiang, Yile Chen, Xiucheng Li, Qin Chao 等ICML 2025
- STEM-LTS: Integrating Semantic-Temporal Dynamics in LLM-driven Time Series AnalysisZhe Zhao, Pengkun Wang, Haibin Wen, Shuang Wang 等AAAI 2025 · 被引用 7 次
- SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding CapabilityJiankang Wang, Zhihan Zhang, Zhihang Liu, Yang Li 等AAAI 2026 · 被引用 20 次
- AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document UnderstandingAhmed Masry, Juan A. Rodríguez, Tianyu Zhang, Suyuchen Wang 等NeurIPS 2025 · 被引用 7 次
