Time-VLM: Exploring Multimodal Vision-Language Models for Augmented Time Series Forecasting
Siru Zhong, Weilin Ruan, Ming Jin, Huan Li, Qingsong Wen, Yuxuan Liang
Abstract
Recent advancements in time series forecasting have explored augmenting models with text or vision modalities to improve accuracy. While text provides contextual understanding, it often lacks fine-grained temporal details. Conversely, vision captures intricate temporal patterns but lacks semantic context, limiting the complementary potential of these modalities. To address this, we propose Time-VLM, a novel multimodal framework that leverages pre-trained Vision-Language Models (VLMs) to bridge temporal, visual, and textual modalities for enhanced forecasting. Our framework comprises three key components: (1) a Retrieval-Augmented Learner, which extracts enriched temporal features through memory bank interactions; (2) a Vision-Augmented Learner, which encodes time series as informative images; and (3) a Text-Augmented Learner, which generates contextual textual descriptions. These components collaborate with frozen pretrained VLMs to produce multimodal embeddings, which are then fused with temporal features for final prediction. Extensive experiments demonstrate that Time-VLM achieves superior performance, particularly in few-shot and zeroshot scenarios, thereby establishing a new direction for multimodal time series forecasting. Code is available at https://github.com/ CityMind-Lab/ICML25-TimeVLM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers16
- Aurora: Towards Universal Generative Multimodal Time Series ForecastingXingjian Wu, Jianxin Jin, Wanghui Qiu, Peng Chen et al.ICLR 2026 · 33 citations
- Can Multimodal LLMs Perform Time Series Anomaly Detection?Xiongxiao Xu, Haoran Wang, Yueqing Liang, Philip S. Yu et al.WWW 2026 · 18 citations
- Multi-Modal View Enhanced Large Vision Models for Long-Term Time Series ForecastingChengAo Shen, Wenchao Yu, Ziming Zhao, Dongjin Song et al.NeurIPS 2025 · 14 citations
- BEDTime: A Unified Benchmark for Automatically Describing Time SeriesMedhasweta Sen, Zachary Gottesman, Jiaxing Qiu, C. Bayan Bruss et al.ICML 2026 · 8 citations
- Unlocking the Value of Text: Event-Driven Reasoning and Multi-Level Alignment for Time Series ForecastingSiyuan Wang, Peng Chen, Yihang Wang, Wanghui Qiu et al.ICLR 2026 · 4 citations
Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang et al.AAAI 2021 · 7,289 citations
- Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series ForecastingHaixu Wu, Jiehui Xu, Jianmin Wang, Mingsheng LongNeurIPS 2021 · 5,824 citations
Related papers
- ST-VLM: A Spatial-to-Image Multimodal Spatial-Temporal Prediction Framework with Vision-Language ModelTong Zhao, Junping Du, Zhe Xue, Meiyu Liang et al.AAAI 2026
- Time-LLM: Time Series Forecasting by Reprogramming Large Language ModelsMing Jin, Shiyu Wang, Lintao Ma, Zhixuan Chu et al.ICLR 2024 · 915 citations
- M3Time: LLM-Enhanced Multi-Modal, Multi-Scale, and Multi-Frequency Multivariate Time Series ForecastingShuning Jia, Baijun Song, Canming Ye, Chun YuanAAAI 2026 · 1 citation
- ChatTime: A Unified Multimodal Time Series Foundation Model Bridging Numerical and Textual DataChengsen Wang, Qi Qi, Jingyu Wang, Haifeng Sun et al.AAAI 2025 · 109 citations
- TimeCMA: Towards LLM-Empowered Multivariate Time Series Forecasting via Cross-Modality AlignmentChenxi Liu, Qianxiong Xu, Hao Miao, Sun Yang et al.AAAI 2025 · 141 citations
