Approaching Human-Level Forecasting with Language Models
Danny Halawi, Fred Zhang, Yueh-Han Chen, Jacob Steinhardt
Abstract
Forecasting future events is important for policy and decision making. In this work, we study whether language models (LMs) can forecast at the level of competitive human forecasters. Towards this goal, we develop a retrieval-augmented LM system designed to automatically search for relevant information, generate forecasts, and aggregate predictions. To facilitate our study, we collect a large dataset of questions from competitive forecasting platforms. Under a test set published after the knowledge cut-offs of our LMs, we evaluate the end-to-end performance of our system against the aggregates of human forecasts. On average, the system nears the crowd aggregate of competitive forecasters, and in some settings surpasses it. Our work suggests that using LMs to forecast the future could provide accurate predictions at scale and help to inform institutional decision making.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 44e5a7d8-0ffe-4057-87f1-e107e2e83a97Cited by top-tier papers19
- LLM-as-a-Prophet: Understanding Predictive Intelligence with Prophet ArenaQingchuan Yang, Simon Mahns, Sida Li, Anri Gu et al.ICLR 2026 · 36 citations
- Argumentative Large Language Models for Explainable and Contestable Claim VerificationGabriel Freedman, Adam Dejl, Deniz Gorur, Xiang Yin et al.AAAI 2025 · 31 citations
- Pitfalls in Evaluating Language Model ForecastersDaniel Paleka, Shashwat Goel, Jonas Geiping, Florian TramèrICLR 2026 · 25 citations
- Predicting Empirical AI Research Outcomes with Language ModelsJiaxin Wen, Chenglei Si, Yueh-Han Chen, He He et al.NeurIPS 2025 · 18 citations
- SimpleStrat: Diversifying Language Model Generation with StratificationJustin Wong, Yury Orlovskiy, Alexander Shypula, Michael Luo et al.NeurIPS 2025 · 17 citations
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Time-LLM: Time Series Forecasting by Reprogramming Large Language ModelsMing Jin, Shiyu Wang, Lintao Ma, Zhixuan Chu et al.ICLR 2024 · 915 citations
- Large Language Models Are Zero-Shot Time Series ForecastersNate Gruver, Marc Finzi, Shikai Qiu, Andrew Gordon WilsonNeurIPS 2023 · 898 citations
- A decoder-only foundation model for time-series forecastingAbhimanyu Das, Weihao Kong, Rajat Sen, Yichen ZhouICML 2024 · 601 citations
Related papers
- Curating the Future: A Scalable Recipe for Training Open-Ended ForecastersNikhil Chandak, Shashwat Goel, Ameya Pandurang Prabhu, Moritz Hardt et al.ICML 2026
- ForecastBench: A Dynamic Benchmark of AI Forecasting CapabilitiesEzra Karger, Houtan Bastani, Yueh-Han Chen, Zachary Jacobs et al.ICLR 2025
- Active Retrieval Augmented GenerationZhengbao Jiang, Frank F. Xu, Luyu Gao, Zhiqing Sun et al.EMNLP 2023 · 315 citations
- Are LLMs Prescient? A Continuous Evaluation using Daily News as the OracleHui Dai, Ryan Teehan, Mengye RenICML 2025
- Augur: Modeling Covariate Causal Associations in Time Series via Large Language ModelsZhiqing Cui, Binwu Wang, Qingxiang Liu, Yeqiang Wang et al.ACL 2026
