Overcoming the Modality Gap in Context-Aided Forecasting
Vincent Zheng, Étienne Marcotte, Arjun Ashok, Andrew Williams, Lijun Sun, Alexandre Drouin, Valentina Zantedeschi
Abstract
Context-aided forecasting (CAF) holds promise for integrating domain knowledge and forward-looking information, enabling AI systems to surpass traditional statistical methods. However, recent empirical studies reveal a puzzling gap: multimodal models often fail to outperform their unimodal counterparts. We hypothesize that this underperformance stems partly from insufficiently verified context usefulness in existing datasets. To address these limitations, we introduce a semi-synthetic data augmentation method that generates contexts both descriptive of temporal dynamics and verifiably complementary to numerical histories. This approach enables massive-scale dataset creation, resulting in CAF-7M, a corpus of 7 million context-augmented time series windows, including a rigorously verified test set. We demonstrate that semi-synthetic pre-training transfers effectively to real-world evaluation, and show clear evidence of context utilization. Our results suggest that dataset quality is a major bottleneck in context-aided forecasting, and that verified context can substantially improve the usefulness of CAF training data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0d37ef99-dd45-456d-8fd0-0a6d81573bd3Builds on12
- Time-LLM: Time Series Forecasting by Reprogramming Large Language ModelsMing Jin, Shiyu Wang, Lintao Ma, Zhixuan Chu et al.ICLR 2024 · 915 citations
- Large Language Models Are Zero-Shot Time Series ForecastersNate Gruver, Marc Finzi, Shikai Qiu, Andrew Gordon WilsonNeurIPS 2023 · 898 citations
- Are Language Models Actually Useful for Time Series Forecasting?Mingtian Tan, Mike A. Merrill, Vinayak Gupta, Tim Althoff et al.NeurIPS 2024 · 326 citations
- UniTime: A Language-Empowered Unified Model for Cross-Domain Time Series ForecastingXu Liu, Junfeng Hu, Yuan Li, Shizhe Diao et al.WWW 2024 · 198 citations
- From News to Forecast: Integrating Event Analysis in LLM-Based Time Series Forecasting with ReflectionXinlei Wang, Maike Feng, Jing Qiu, Jinjin Gu et al.NeurIPS 2024 · 181 citations
Related papers
- Rethinking Multimodal Time-Series Forecasting EvaluationHaoxin Liu, Yichen Zhou, Rajat Sen, B. Aditya Prakash et al.ICML 2026 · 3 citations
- SK-VQA: Synthetic Knowledge Generation at Scale for Training Context-Augmented Multimodal LLMsXin Su, Man Luo, Kris W. Pan, Tien Pei Chou et al.ICML 2025
- BLAST: Balanced Sampling Time Series Corpus for Universal Forecasting ModelsZezhi Shao, Yujie Li, Fei Wang, Chengqing Yu et al.KDD 2025 · 3 citations
- Context is Key: A Benchmark for Forecasting with Essential Textual InformationAndrew Robert Williams, Arjun Ashok, Étienne Marcotte, Valentina Zantedeschi et al.ICML 2025
- Multi-view Self-Supervised Contrastive Learning for Multivariate Time SeriesYuhan Wu, Xiyu Meng, Yang He, Junru Zhang et al.ACM MM 2024 · 5 citations
