Consistency Checks for Language Model Forecasters
Daniel Paleka, Abhimanyu Pallavi Sudhir, Alejandro Alvarez, Vineeth Bhat, Adam Shen, Evan Wang, Florian Tramèr
摘要
Forecasting is a task that is difficult to evaluate: the ground truth can only be known in the future. Recent work showing LLM forecasters rapidly approaching human-level performance begs the question: how can we benchmark and evaluate these forecasters instantaneously? Following the consistency check framework, we measure the performance of forecasters in terms of the consistency of their predictions on different logically-related questions. We propose a new, general consistency metric based on arbitrage: for example, if a forecasting AI illogically predicts that both the Democratic and Republican parties have 60% probability of winning the 2024 US presidential election, an arbitrageur can trade against the forecaster's predictions and make a profit. We build an automated evaluation system that generates a set of base questions, instantiates consistency checks from these questions, elicits the predictions of the forecaster, and measures the consistency of the predictions. We then build a standard, proper-scoring-rule forecasting benchmark, and show that our (instantaneous) consistency metrics correlate with LLM forecasters' ground truth Brier scores (which are only known in the future). We also release a consistency benchmark that resolves in 2028, providing a long-term evaluation tool for forecasting.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- LLM-as-a-Prophet: Understanding Predictive Intelligence with Prophet ArenaQingchuan Yang, Simon Mahns, Sida Li, Anri Gu 等ICLR 2026 · 被引用 36 次
- Pitfalls in Evaluating Language Model ForecastersDaniel Paleka, Shashwat Goel, Jonas Geiping, Florian TramèrICLR 2026 · 被引用 25 次
- Analytica: Soft Propositional Reasoning for Robust and Scalable LLM-Driven AnalysisJunyan Cheng, Kyle Richardson, Peter ChinICLR 2026 · 被引用 4 次
- Adaptively profiling models with task elicitationDavis Brown, Prithvi Balehannina, Helen Jin, Shreya Havaldar 等EMNLP 2025 · 被引用 1 次
- LogiConBench: Benchmarking Logical Consistencies of LLMsZheng Chen, Chuan Zhou, Fengxiang Cheng, Tin Po Yip 等ICLR 2026
它引用的顶会 Paper8
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Approaching Human-Level Forecasting with Language ModelsDanny Halawi, Fred Zhang, Yueh-Han Chen, Jacob SteinhardtNeurIPS 2024 · 被引用 142 次
- Discovering Latent Knowledge in Language Models Without SupervisionCollin Burns, Haotian Ye, Dan Klein, Jacob SteinhardtICLR 2023 · 被引用 45 次
- Benchmarking and Improving Generator-Validator Consistency of Language ModelsXiang Lisa Li, Vaishnavi Shrivastava, Siyan Li, Tatsunori Hashimoto 等ICLR 2024 · 被引用 45 次
- Goodhart's Law in Reinforcement LearningJacek Karwowski, Oliver Hayman, Xingjian Bai, Klaus Kiendlhofer 等ICLR 2024 · 被引用 22 次
相关 Paper
- ForecastBench: A Dynamic Benchmark of AI Forecasting CapabilitiesEzra Karger, Houtan Bastani, Yueh-Han Chen, Zachary Jacobs 等ICLR 2025
- Martingale Score: An Unsupervised Metric for Bayesian Rationality in LLM ReasoningZhonghao He, Tianyi Alex Qiu, Hirokazu Shirado, Maarten SapNeurIPS 2025 · 被引用 6 次
- Measuring Consistency in Text-based Financial Forecasting ModelsLinyi Yang, Yingpeng Ma, Yue ZhangACL 2023 · 被引用 1 次
- Benchmarking at the Edge of ComprehensionSamuele Marro, Jialin Yu, Emanuele La Malfa, Oishi Deb 等ICML 2026
- Are LLMs Prescient? A Continuous Evaluation using Daily News as the OracleHui Dai, Ryan Teehan, Mengye RenICML 2025
