Pitfalls in Evaluating Language Model Forecasters
Daniel Paleka, Shashwat Goel, Jonas Geiping, Florian Tramèr
Abstract
Large language models (LLMs) have recently been applied to forecasting tasks, with some works claiming these systems match or exceed human performance. In this paper, we argue that, as a community, we should be careful about such conclusions as evaluating LLM forecasters presents unique challenges. We identify two broad categories of issues: (1) difficulty in trusting evaluation results due to many forms of temporal leakage, and (2) difficulty in extrapolating from evaluation performance to real-world forecasting. Through systematic analysis and concrete examples from prior work, we demonstrate how evaluation flaws can raise concerns about current and future performance claims. We argue that more rigorous evaluation methodologies are needed to confidently assess the forecasting abilities of LLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 91f368b5-34c4-42ba-9e6b-4b3731e345aaCited by top-tier papers3
- FutureX: An Advanced Live Benchmark for LLM Agents in Future PredictionZhiyuan Zeng, Jiashuo Liu, Siyuan Chen, Tianci He et al.ICLR 2026 · 51 citations
- Curating the Future: A Scalable Recipe for Training Open-Ended ForecastersNikhil Chandak, Shashwat Goel, Ameya Pandurang Prabhu, Moritz Hardt et al.ICML 2026
- Global Merger-Arbitrage Forecasting with Language ModelsHinal Jajal, Michał Mucha, Charles Sweat, Chris Pulman et al.ICML 2026
Builds on3
- Approaching Human-Level Forecasting with Language ModelsDanny Halawi, Fred Zhang, Yueh-Han Chen, Jacob SteinhardtNeurIPS 2024 · 142 citations
- Consistency Checks for Language Model ForecastersDaniel Paleka, Abhimanyu Pallavi Sudhir, Alejandro Alvarez, Vineeth Bhat et al.ICLR 2025
- Are LLMs Prescient? A Continuous Evaluation using Daily News as the OracleHui Dai, Ryan Teehan, Mengye RenICML 2025
Related papers
- A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and RecommendationsMd. Tahmid Rahman Laskar, Sawsan Alqahtani, M. Saiful Bari, Mizanur Rahman et al.EMNLP 2024 · 47 citations
- LLM-as-a-Prophet: Understanding Predictive Intelligence with Prophet ArenaQingchuan Yang, Simon Mahns, Sida Li, Anri Gu et al.ICLR 2026 · 36 citations
- Are Language Models Actually Useful for Time Series Forecasting?Mingtian Tan, Mike A. Merrill, Vinayak Gupta, Tim Althoff et al.NeurIPS 2024 · 326 citations
- ForecastBench: A Dynamic Benchmark of AI Forecasting CapabilitiesEzra Karger, Houtan Bastani, Yueh-Han Chen, Zachary Jacobs et al.ICLR 2025
- LH-DECEPTION: Simulating and Understanding LLM Deceptive Behaviors in Long-Horizon InteractionsYang Xu, Xuanming Zhang, Samuel (Min-Hsuan) Yeh, Jwala Dhamala et al.ICLR 2026 · 7 citations
