Approaching Human-Level Forecasting with Language Models
Danny Halawi, Fred Zhang, Yueh-Han Chen, Jacob Steinhardt
摘要
Forecasting future events is important for policy and decision making. In this work, we study whether language models (LMs) can forecast at the level of competitive human forecasters. Towards this goal, we develop a retrieval-augmented LM system designed to automatically search for relevant information, generate forecasts, and aggregate predictions. To facilitate our study, we collect a large dataset of questions from competitive forecasting platforms. Under a test set published after the knowledge cut-offs of our LMs, we evaluate the end-to-end performance of our system against the aggregates of human forecasts. On average, the system nears the crowd aggregate of competitive forecasters, and in some settings surpasses it. Our work suggests that using LMs to forecast the future could provide accurate predictions at scale and help to inform institutional decision making.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- LLM-as-a-Prophet: Understanding Predictive Intelligence with Prophet ArenaQingchuan Yang, Simon Mahns, Sida Li, Anri Gu 等ICLR 2026 · 被引用 36 次
- Argumentative Large Language Models for Explainable and Contestable Claim VerificationGabriel Freedman, Adam Dejl, Deniz Gorur, Xiang Yin 等AAAI 2025 · 被引用 31 次
- Pitfalls in Evaluating Language Model ForecastersDaniel Paleka, Shashwat Goel, Jonas Geiping, Florian TramèrICLR 2026 · 被引用 25 次
- Predicting Empirical AI Research Outcomes with Language ModelsJiaxin Wen, Chenglei Si, Yueh-Han Chen, He He 等NeurIPS 2025 · 被引用 18 次
- SimpleStrat: Diversifying Language Model Generation with StratificationJustin Wong, Yury Orlovskiy, Alexander Shypula, Michael Luo 等NeurIPS 2025 · 被引用 17 次
它引用的顶会 Paper10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Time-LLM: Time Series Forecasting by Reprogramming Large Language ModelsMing Jin, Shiyu Wang, Lintao Ma, Zhixuan Chu 等ICLR 2024 · 被引用 915 次
- Large Language Models Are Zero-Shot Time Series ForecastersNate Gruver, Marc Finzi, Shikai Qiu, Andrew Gordon WilsonNeurIPS 2023 · 被引用 898 次
- A decoder-only foundation model for time-series forecastingAbhimanyu Das, Weihao Kong, Rajat Sen, Yichen ZhouICML 2024 · 被引用 601 次
相关 Paper
- Curating the Future: A Scalable Recipe for Training Open-Ended ForecastersNikhil Chandak, Shashwat Goel, Ameya Pandurang Prabhu, Moritz Hardt 等ICML 2026
- ForecastBench: A Dynamic Benchmark of AI Forecasting CapabilitiesEzra Karger, Houtan Bastani, Yueh-Han Chen, Zachary Jacobs 等ICLR 2025
- Active Retrieval Augmented GenerationZhengbao Jiang, Frank F. Xu, Luyu Gao, Zhiqing Sun 等EMNLP 2023 · 被引用 315 次
- Are LLMs Prescient? A Continuous Evaluation using Daily News as the OracleHui Dai, Ryan Teehan, Mengye RenICML 2025
- Augur: Modeling Covariate Causal Associations in Time Series via Large Language ModelsZhiqing Cui, Binwu Wang, Qingxiang Liu, Yeqiang Wang 等ACL 2026
