Time-VLM: Exploring Multimodal Vision-Language Models for Augmented Time Series Forecasting
Siru Zhong, Weilin Ruan, Ming Jin, Huan Li, Qingsong Wen, Yuxuan Liang
摘要
Recent advancements in time series forecasting have explored augmenting models with text or vision modalities to improve accuracy. While text provides contextual understanding, it often lacks fine-grained temporal details. Conversely, vision captures intricate temporal patterns but lacks semantic context, limiting the complementary potential of these modalities. To address this, we propose Time-VLM, a novel multimodal framework that leverages pre-trained Vision-Language Models (VLMs) to bridge temporal, visual, and textual modalities for enhanced forecasting. Our framework comprises three key components: (1) a Retrieval-Augmented Learner, which extracts enriched temporal features through memory bank interactions; (2) a Vision-Augmented Learner, which encodes time series as informative images; and (3) a Text-Augmented Learner, which generates contextual textual descriptions. These components collaborate with frozen pretrained VLMs to produce multimodal embeddings, which are then fused with temporal features for final prediction. Extensive experiments demonstrate that Time-VLM achieves superior performance, particularly in few-shot and zeroshot scenarios, thereby establishing a new direction for multimodal time series forecasting. Code is available at https://github.com/ CityMind-Lab/ICML25-TimeVLM .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Aurora: Towards Universal Generative Multimodal Time Series ForecastingXingjian Wu, Jianxin Jin, Wanghui Qiu, Peng Chen 等ICLR 2026 · 被引用 33 次
- Can Multimodal LLMs Perform Time Series Anomaly Detection?Xiongxiao Xu, Haoran Wang, Yueqing Liang, Philip S. Yu 等WWW 2026 · 被引用 18 次
- Multi-Modal View Enhanced Large Vision Models for Long-Term Time Series ForecastingChengAo Shen, Wenchao Yu, Ziming Zhao, Dongjin Song 等NeurIPS 2025 · 被引用 14 次
- BEDTime: A Unified Benchmark for Automatically Describing Time SeriesMedhasweta Sen, Zachary Gottesman, Jiaxing Qiu, C. Bayan Bruss 等ICML 2026 · 被引用 8 次
- Unlocking the Value of Text: Event-Driven Reasoning and Multi-Level Alignment for Time Series ForecastingSiyuan Wang, Peng Chen, Yihang Wang, Wanghui Qiu 等ICLR 2026 · 被引用 4 次
它引用的顶会 Paper21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang 等AAAI 2021 · 被引用 7,289 次
- Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series ForecastingHaixu Wu, Jiehui Xu, Jianmin Wang, Mingsheng LongNeurIPS 2021 · 被引用 5,824 次
相关 Paper
- ST-VLM: A Spatial-to-Image Multimodal Spatial-Temporal Prediction Framework with Vision-Language ModelTong Zhao, Junping Du, Zhe Xue, Meiyu Liang 等AAAI 2026
- Time-LLM: Time Series Forecasting by Reprogramming Large Language ModelsMing Jin, Shiyu Wang, Lintao Ma, Zhixuan Chu 等ICLR 2024 · 被引用 915 次
- M3Time: LLM-Enhanced Multi-Modal, Multi-Scale, and Multi-Frequency Multivariate Time Series ForecastingShuning Jia, Baijun Song, Canming Ye, Chun YuanAAAI 2026 · 被引用 1 次
- ChatTime: A Unified Multimodal Time Series Foundation Model Bridging Numerical and Textual DataChengsen Wang, Qi Qi, Jingyu Wang, Haifeng Sun 等AAAI 2025 · 被引用 109 次
- TimeCMA: Towards LLM-Empowered Multivariate Time Series Forecasting via Cross-Modality AlignmentChenxi Liu, Qianxiong Xu, Hao Miao, Sun Yang 等AAAI 2025 · 被引用 141 次
