Curating the Future: A Scalable Recipe for Training Open-Ended Forecasters
Nikhil Chandak, Shashwat Goel, Ameya Pandurang Prabhu, Moritz Hardt, Jonas Geiping
Abstract
High-stakes decision making involves reasoning under uncertainty about the future. In this work, we train language models to make predictions on open-ended forecasting questions. To scale up training data, we synthesize novel forecasting questions from global events reported in daily news. While directly training on this data leads to performance drops, carefully curating questions creates a valuable training resource. We use the resulting dataset, OpenForesight, to post-train Qwen3 thinking models. To prevent leakage of future information during training and evaluation, we use an offline news corpus, both for data generation and retrieval in our forecasting system. Guided by a small validation set, we show the benefits of retrieval, and an improved reward function for reinforcement learning (RL). Once we obtain our final forecasting system, we perform held-out testing between May to August 2025. Our specialized model, OpenForecaster-8B, matches much larger proprietary models, with our training improving the accuracy, calibration, and consistency of predictions. We find calibration improvements from forecasting training generalize across popular benchmarks. We will open-source our models, code, and data to make LLM based forecasting research broadly accessible.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 35b890b6-03f9-4fff-9727-2e249a7b8284Builds on8
- Approaching Human-Level Forecasting with Language ModelsDanny Halawi, Fred Zhang, Yueh-Han Chen, Jacob SteinhardtNeurIPS 2024 · 142 citations
- Beyond Binary Rewards: Training LMs to Reason About Their UncertaintyMehul Damani, Isha Puri, Stewart Slocum, Idan Shenfeld et al.ICLR 2026 · 116 citations
- FutureX: An Advanced Live Benchmark for LLM Agents in Future PredictionZhiyuan Zeng, Jiashuo Liu, Siyuan Chen, Tianci He et al.ICLR 2026 · 51 citations
- Pitfalls in Evaluating Language Model ForecastersDaniel Paleka, Shashwat Goel, Jonas Geiping, Florian TramèrICLR 2026 · 25 citations
- Consistency Checks for Language Model ForecastersDaniel Paleka, Abhimanyu Pallavi Sudhir, Alejandro Alvarez, Vineeth Bhat et al.ICLR 2025
Related papers
- ReSearch: Learning to Reason with Search for LLMs via Reinforcement LearningMingyang Chen, Linzhuang Sun, Tianpeng Li, Haoze Sun et al.NeurIPS 2025 · 125 citations
- Are LLMs Prescient? A Continuous Evaluation using Daily News as the OracleHui Dai, Ryan Teehan, Mengye RenICML 2025
- OpenEstimate: Evaluating LLMs on Reasoning Under Uncertainty with Real-World DataAlana Renda, Jillian Ross, Jacob AndreasICLR 2026 · 3 citations
- Train for Truth, Keep the Skills: Binary Retrieval-Augmented Reward Mitigates HallucinationsTong Chen, Akari Asai, Luke Zettlemoyer, Hannaneh Hajishirzi et al.ICML 2026
- ForecastBench: A Dynamic Benchmark of AI Forecasting CapabilitiesEzra Karger, Houtan Bastani, Yueh-Han Chen, Zachary Jacobs et al.ICLR 2025
