Consistency Checks for Language Model Forecasters
Daniel Paleka, Abhimanyu Pallavi Sudhir, Alejandro Alvarez, Vineeth Bhat, Adam Shen, Evan Wang, Florian Tramèr
Abstract
Forecasting is a task that is difficult to evaluate: the ground truth can only be known in the future. Recent work showing LLM forecasters rapidly approaching human-level performance begs the question: how can we benchmark and evaluate these forecasters instantaneously? Following the consistency check framework, we measure the performance of forecasters in terms of the consistency of their predictions on different logically-related questions. We propose a new, general consistency metric based on arbitrage: for example, if a forecasting AI illogically predicts that both the Democratic and Republican parties have 60% probability of winning the 2024 US presidential election, an arbitrageur can trade against the forecaster's predictions and make a profit. We build an automated evaluation system that generates a set of base questions, instantiates consistency checks from these questions, elicits the predictions of the forecaster, and measures the consistency of the predictions. We then build a standard, proper-scoring-rule forecasting benchmark, and show that our (instantaneous) consistency metrics correlate with LLM forecasters' ground truth Brier scores (which are only known in the future). We also release a consistency benchmark that resolves in 2028, providing a long-term evaluation tool for forecasting.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 937af11f-7b40-4f57-8dc7-acf3dd03fbc6Cited by top-tier papers7
- LLM-as-a-Prophet: Understanding Predictive Intelligence with Prophet ArenaQingchuan Yang, Simon Mahns, Sida Li, Anri Gu et al.ICLR 2026 · 36 citations
- Pitfalls in Evaluating Language Model ForecastersDaniel Paleka, Shashwat Goel, Jonas Geiping, Florian TramèrICLR 2026 · 25 citations
- Analytica: Soft Propositional Reasoning for Robust and Scalable LLM-Driven AnalysisJunyan Cheng, Kyle Richardson, Peter ChinICLR 2026 · 4 citations
- Adaptively profiling models with task elicitationDavis Brown, Prithvi Balehannina, Helen Jin, Shreya Havaldar et al.EMNLP 2025 · 1 citation
- LogiConBench: Benchmarking Logical Consistencies of LLMsZheng Chen, Chuan Zhou, Fengxiang Cheng, Tin Po Yip et al.ICLR 2026
Builds on8
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Approaching Human-Level Forecasting with Language ModelsDanny Halawi, Fred Zhang, Yueh-Han Chen, Jacob SteinhardtNeurIPS 2024 · 142 citations
- Discovering Latent Knowledge in Language Models Without SupervisionCollin Burns, Haotian Ye, Dan Klein, Jacob SteinhardtICLR 2023 · 45 citations
- Benchmarking and Improving Generator-Validator Consistency of Language ModelsXiang Lisa Li, Vaishnavi Shrivastava, Siyan Li, Tatsunori Hashimoto et al.ICLR 2024 · 45 citations
- Goodhart's Law in Reinforcement LearningJacek Karwowski, Oliver Hayman, Xingjian Bai, Klaus Kiendlhofer et al.ICLR 2024 · 22 citations
Related papers
- ForecastBench: A Dynamic Benchmark of AI Forecasting CapabilitiesEzra Karger, Houtan Bastani, Yueh-Han Chen, Zachary Jacobs et al.ICLR 2025
- Martingale Score: An Unsupervised Metric for Bayesian Rationality in LLM ReasoningZhonghao He, Tianyi Alex Qiu, Hirokazu Shirado, Maarten SapNeurIPS 2025 · 6 citations
- Measuring Consistency in Text-based Financial Forecasting ModelsLinyi Yang, Yingpeng Ma, Yue ZhangACL 2023 · 1 citation
- Benchmarking at the Edge of ComprehensionSamuele Marro, Jialin Yu, Emanuele La Malfa, Oishi Deb et al.ICML 2026
- Are LLMs Prescient? A Continuous Evaluation using Daily News as the OracleHui Dai, Ryan Teehan, Mengye RenICML 2025
