ForecastBench: A Dynamic Benchmark of AI Forecasting Capabilities
Ezra Karger, Houtan Bastani, Yueh-Han Chen, Zachary Jacobs, Danny Halawi, Fred Zhang, Philip Tetlock
摘要
Forecasts of future events are essential inputs into informed decision-making. Machine learning (ML) systems have the potential to deliver forecasts at scale, but there is no framework for evaluating the accuracy of ML systems on a standardized set of forecasting questions. To address this gap, we introduce ForecastBench: a dynamic benchmark that evaluates the accuracy of ML systems on an automatically generated and regularly updated set of 1,000 forecasting questions. To avoid any possibility of data leakage, ForecastBench is comprised solely of questions about future events that have no known answer at the time of submission. We quantify the capabilities of current ML systems by collecting forecasts from expert (human) forecasters, the general public, and LLMs on a random subset of questions from the benchmark (N = 200). While LLMs have achieved super-human performance on many benchmarks, they perform less well here: expert forecasters outperform the top-performing LLM (p-value < 0.001). We display system and human scores in a public leaderboard at www.forecastbench.org.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- FutureX: An Advanced Live Benchmark for LLM Agents in Future PredictionZhiyuan Zeng, Jiashuo Liu, Siyuan Chen, Tianci He 等ICLR 2026 · 被引用 51 次
- LLM-as-a-Prophet: Understanding Predictive Intelligence with Prophet ArenaQingchuan Yang, Simon Mahns, Sida Li, Anri Gu 等ICLR 2026 · 被引用 36 次
- Predicting Empirical AI Research Outcomes with Language ModelsJiaxin Wen, Chenglei Si, Yueh-Han Chen, He He 等NeurIPS 2025 · 被引用 18 次
- Inferring Events from Time Series using Language ModelsMingtian Tan, Mike A. Merrill, Zachary Gottesman, Tim Althoff 等ACL 2026 · 被引用 7 次
- FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMsQian Chen, Jinlan Fu, Changsong Li, Min zhang 等ICML 2026 · 被引用 5 次
它引用的顶会 Paper13
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos 等ICML 2024 · 被引用 1,212 次
- Time-LLM: Time Series Forecasting by Reprogramming Large Language ModelsMing Jin, Shiyu Wang, Lintao Ma, Zhixuan Chu 等ICLR 2024 · 被引用 915 次
- Large Language Models Are Zero-Shot Time Series ForecastersNate Gruver, Marc Finzi, Shikai Qiu, Andrew Gordon WilsonNeurIPS 2023 · 被引用 898 次
- A decoder-only foundation model for time-series forecastingAbhimanyu Das, Weihao Kong, Rajat Sen, Yichen ZhouICML 2024 · 被引用 601 次
相关 Paper
- Approaching Human-Level Forecasting with Language ModelsDanny Halawi, Fred Zhang, Yueh-Han Chen, Jacob SteinhardtNeurIPS 2024 · 被引用 142 次
- LiveBench: A Challenging, Contamination-Limited LLM BenchmarkColin White, Samuel Dooley, Manley Roberts, Arka Pal 等ICLR 2025
- TFRBench: A Reasoning Benchmark for Evaluating Forecasting SystemsMd Atik Ahamed, Mihir Parmar, Palash Goyal, Yiwen Song 等ICML 2026
- Are LLMs Prescient? A Continuous Evaluation using Daily News as the OracleHui Dai, Ryan Teehan, Mengye RenICML 2025
- Pitfalls in Evaluating Language Model ForecastersDaniel Paleka, Shashwat Goel, Jonas Geiping, Florian TramèrICLR 2026 · 被引用 25 次
