MM-Forecast: A Multimodal Approach to Temporal Event Forecasting with Large Language Models
Haoxuan Li, Zhengmao Yang, Yunshan Ma, Yi Bin, Yang Yang, Tat-Seng Chua
Abstract
We study an emerging and intriguing problem of multimodal temporal event forecasting with large language models. Compared to using text or graph modalities, the investigation of utilizing images for temporal event forecasting has not been fully explored, especially in the era of large language models (LLMs). To bridge this gap, we are particularly interested in two key questions of: 1) why images will help in temporal event forecasting, and 2) how to integrate images into the LLM-based forecasting framework. To answer these research questions, we propose to identify two essential functions that images play in the scenario of temporal event forecasting, i.e., highlighting and complementary. Then, we develop a novel framework, named MM-Forecast. It employs an Image Function Identification module to recognize these functions as verbal descriptions using multimodal large language models (MLLMs), and subsequently incorporates these function descriptions into LLM-based forecasting models. To evaluate our approach, we construct a new multimodal dataset, MidEast-TE-mm, by extending an existing event dataset MidEast-TE-mini with images. Empirical studies demonstrate that our MM-Forecast can correctly identify the image functions, and further more, incorporating these verbal function descriptions significantly improves the forecasting performance. The dataset, code, and prompts are available at https://github.com/LuminosityX/MM-Forecast.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 052002eb-2985-4ec9-8d25-957e758ad516Cited by top-tier papers3
- DAMMFND: Domain-Aware Multimodal Multi-view Fake News DetectionWeihai Lu, Yu Tong, Zhiqiu YeAAAI 2025 · 21 citations
- FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMsQian Chen, Jinlan Fu, Changsong Li, Min zhang et al.ICML 2026 · 5 citations
- TimeSoccer: An End-to-End Multimodal Large Language Model for Soccer Commentary GenerationLing You, Wenxuan Huang, Xinni Xie, Xiangyi Wei et al.ACM MM 2025 · 2 citations
Builds on17
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- Recurrent Event Network: Autoregressive Structure Inferenceover Temporal Knowledge GraphsWoojeong Jin, Meng Qu, Xisen Jin, Xiang RenEMNLP 2020 · 353 citations
Related papers
- M3Time: LLM-Enhanced Multi-Modal, Multi-Scale, and Multi-Frequency Multivariate Time Series ForecastingShuning Jia, Baijun Song, Canming Ye, Chun YuanAAAI 2026 · 1 citation
- Multi-Modal View Enhanced Large Vision Models for Long-Term Time Series ForecastingChengAo Shen, Wenchao Yu, Ziming Zhao, Dongjin Song et al.NeurIPS 2025 · 14 citations
- CLLMate: A Multimodal Benchmark for Weather and Climate Events ForecastingHaobo Li, Zhaowei Wang, Jiachen Wang, Yueya Wang et al.EMNLP 2025 · 2 citations
- DualSG: A Dual-Stream Explicit Semantic-Guided Multivariate Time Series Forecasting FrameworkKuiye Ding, Fanda Fan, Yao Wang, Ruijie Jian et al.ACM MM 2025 · 6 citations
- TimeCAP: Learning to Contextualize, Augment, and Predict Time Series Events with Large Language Model AgentsGeon Lee, Wenchao Yu, Kijung Shin, Wei Cheng et al.AAAI 2025 · 39 citations
