Context is Key: A Benchmark for Forecasting with Essential Textual Information
Andrew Robert Williams, Arjun Ashok, Étienne Marcotte, Valentina Zantedeschi, Jithendaraa Subramanian, Roland Riachi, James Requeima, Alexandre Lacoste, Irina Rish, Nicolas Chapados, Alexandre Drouin
Abstract
Forecasting is a critical task in decision making across various domains. While numerical data provides a foundation, it often lacks crucial context necessary for accurate predictions. Human forecasters frequently rely on additional information, such as background knowledge or constraints, which can be efficiently communicated through natural language. However, the ability of existing forecasting models to effectively integrate this textual information remains an open question. To address this, we introduce "Context is Key" (CiK), a time series forecasting benchmark that pairs numerical data with diverse types of carefully crafted textual context, requiring models to integrate both modalities. We evaluate a range of approaches, including statistical models, time series foundation models, and LLMbased forecasters, and propose a simple yet effective LLM prompting method that outperforms all other tested methods on our benchmark. Our experiments highlight the importance of incorporating contextual information, demonstrate surprising performance when using LLM-based forecasting models, and also reveal some of their critical shortcomings. By presenting this benchmark, we aim to advance multimodal forecasting, promoting models that are both accurate and accessible to decision-makers with varied technical expertise. The benchmark can be visualized at https://servicenow.github.io/context-is-key-forecasting/v0/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e40f37d6-bc84-4de1-bff3-f38c3f6a06eaCited by top-tier papers15
- Language in the Flow of Time: Time-Series-Paired Texts Weaved into a Unified Temporal NarrativeZihao Li, Xiao Lin, Zhining Liu, Jiaru Zou et al.ICLR 2026 · 41 citations
- True Zero-Shot Inference of Dynamical Systems Preserving Long-Term StatisticsChristoph Jürgen Hemmer, Daniel DurstewitzNeurIPS 2025 · 25 citations
- Improving Time Series Forecasting via Instance-aware Post-hoc RevisionZhiding Liu, Mingyue Cheng, Guanhao Zhao, Jiqian Yang et al.NeurIPS 2025 · 16 citations
- Inferring Events from Time Series using Language ModelsMingtian Tan, Mike A. Merrill, Zachary Gottesman, Tim Althoff et al.ACL 2026 · 7 citations
- TimeRecipe: A Time-Series Forecasting Recipe via Benchmarking Module Level EffectivenessZhiyuan Zhao, Juntong Ni, Shangqing Xu, Haoxin Liu et al.ICLR 2026 · 7 citations
Builds on9
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang et al.AAAI 2021 · 7,289 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Time-LLM: Time Series Forecasting by Reprogramming Large Language ModelsMing Jin, Shiyu Wang, Lintao Ma, Zhixuan Chu et al.ICLR 2024 · 915 citations
- Large Language Models Are Zero-Shot Time Series ForecastersNate Gruver, Marc Finzi, Shikai Qiu, Andrew Gordon WilsonNeurIPS 2023 · 898 citations
- Unified Training of Universal Time Series Forecasting TransformersGerald Woo, Chenghao Liu, Akshat Kumar, Caiming Xiong et al.ICML 2024 · 513 citations
Related papers
- Rethinking Multimodal Time-Series Forecasting EvaluationHaoxin Liu, Yichen Zhou, Rajat Sen, B. Aditya Prakash et al.ICML 2026 · 3 citations
- M3Time: LLM-Enhanced Multi-Modal, Multi-Scale, and Multi-Frequency Multivariate Time Series ForecastingShuning Jia, Baijun Song, Canming Ye, Chun YuanAAAI 2026 · 1 citation
- Temporal Knowledge Graph Forecasting Without Knowledge Using In-Context LearningDong-Ho Lee, Kian Ahrabian, Woojeong Jin, Fred Morstatter et al.EMNLP 2023 · 30 citations
- TFRBench: A Reasoning Benchmark for Evaluating Forecasting SystemsMd Atik Ahamed, Mihir Parmar, Palash Goyal, Yiwen Song et al.ICML 2026
- ExoTimer: Leveraging Large Language Models for Time Series Forecasting with Exogenous VariablesLan Wu, Xuebin Wang, Chenglong Ge, Ruijuan Chu et al.AAAI 2026
