TFRBench: A Reasoning Benchmark for Evaluating Forecasting Systems
Md Atik Ahamed, Mihir Parmar, Palash Goyal, Yiwen Song, Long T. Le, Qiang Cheng, Chun-Liang Li, Hamid Palangi, Jinsung Yoon, Tomas Pfister
Abstract
We introduce TFRBench, the first benchmark designed to evaluate the reasoning capabilities of forecasting systems. Traditionally, time-series forecasting has been evaluated solely on numerical accuracy, treating foundation models as "black boxes." Unlike existing benchmarks, TFRBench provides a protocol for evaluating the reasoning generated by forecasting systems--specifically their analysis of cross-channel dependencies, trends, and external events. To enable this, we propose a systematic multi-agent framework that utilizes an iterative verification loop to synthesize numerically grounded reasoning traces. Spanning ten datasets across five domains, our evaluation confirms that this reasoning is causally effective; useful for evaluation; and prompting LLMs with our generated traces significantly improves forecasting accuracy compared to direct numerical prediction (e.g., avg. % %), validating the quality of our reasoning. Conversely, benchmarking experiments reveal that off-the-shelf LLMs consistently struggle with both reasoning (lower LLM-as-a-Judge scores) and numerical forecasting, frequently failing to capture domain-specific dynamics. TFRBench thus establishes a new standard for interpretable, reasoning-based evaluation in time-series forecasting. Our benchmark is available at: https://tfrbench.github.io
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1f226249-8737-4c8d-abb5-2a98d9b1b051Builds on4
- iTransformer: Inverted Transformers Are Effective for Time Series ForecastingYong Liu, Tengge Hu, Haoran Zhang, Haixu Wu et al.ICLR 2024 · 1,703 citations
- Large Language Models Are Zero-Shot Time Series ForecastersNate Gruver, Marc Finzi, Shikai Qiu, Andrew Gordon WilsonNeurIPS 2023 · 898 citations
- A decoder-only foundation model for time-series forecastingAbhimanyu Das, Weihao Kong, Rajat Sen, Yichen ZhouICML 2024 · 601 citations
- INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based AgentHaohang Li, Yupeng Cao, Yangyang Yu, Shashidhar Reddy Javaji et al.ACL 2025
Related papers
- TimeSeriesExamAgent: Creating Time Series Reasoning Benchmarks at ScaleMalgorzata Gwiazda, Yifu Cai, Mononito Goswami, Arjun Choudhry et al.ICLR 2026 · 6 citations
- TSRBench: A Comprehensive Multi-task Multi-modal Time Series Reasoning Benchmark for Generalist ModelsFangxu Yu, Xingang Guo, Lingzhi Yuan, Haoqiang Kang et al.ICML 2026
- ForecastBench: A Dynamic Benchmark of AI Forecasting CapabilitiesEzra Karger, Houtan Bastani, Yueh-Han Chen, Zachary Jacobs et al.ICLR 2025
- A²RBench: An Automatic Paradigm for Formally Verifiable Abstract Reasoning Benchmark GenerationQingchuan Ma, Yuexiao Ma, Yongkang Xie, Tianyu Xie et al.ICML 2026 · 1 citation
- SPAN: Benchmarking and Improving Cross-Calendar Temporal Reasoning of Large Language ModelsZhongjian Miao, Hao Fu, Chen WeiAAAI 2026
