FutureX: An Advanced Live Benchmark for LLM Agents in Future Prediction
Zhiyuan Zeng, Jiashuo Liu, Siyuan Chen, Tianci He, Yali Liao, Yixiao Tian, Jinpeng Wang, Zaiyuan Wang, Yang Yang, Lingyue Yin, Mingren Yin, Zhenwei Zhu
Abstract
Future prediction is a complex task for LLM agents, requiring a high level of analytical thinking, information gathering, contextual understanding, and decision-making under uncertainty. Agents must not only gather and interpret vast amounts of dynamic information but also integrate diverse data sources, weigh uncertainties, and adapt predictions based on emerging trends, just as human experts do in fields like politics, economics, and finance. Despite its importance, no large-scale benchmark exists for evaluating agents on future prediction, largely due to challenges in handling real-time updates and retrieving timely, accurate answers. To address this, we introduce , a dynamic and live evaluation benchmark specifically designed for LLM agents performing future prediction tasks. FutureX is the largest and most diverse live benchmark for future prediction, supporting real-time daily updates and eliminating data contamination through an automated pipeline for question gathering and answer collection. We evaluate 25 LLM/agent models, including those with reasoning, search capabilities, and integration of external tools such as the open-source Deep Research Agent and closed-source Deep Research models. This comprehensive evaluation assesses agents'adaptive reasoning and performance in dynamic environments. Additionally, we provide in-depth analyses of agents'failure modes and performance pitfalls in future-oriented tasks, including the vulnerability to fake web pages and the temporal validity. Our goal is to establish a dynamic, contamination-free evaluation standard that drives the development of LLM agents capable of performing at the level of professional human analysts in complex reasoning and predictive thinking.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fc1da945-239b-4ea1-9372-477157069c31Cited by top-tier papers7
- LLM-as-a-Prophet: Understanding Predictive Intelligence with Prophet ArenaQingchuan Yang, Simon Mahns, Sida Li, Anri Gu et al.ICLR 2026 · 36 citations
- FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMsQian Chen, Jinlan Fu, Changsong Li, Min zhang et al.ICML 2026 · 5 citations
- Analytica: Soft Propositional Reasoning for Robust and Scalable LLM-Driven AnalysisJunyan Cheng, Kyle Richardson, Peter ChinICLR 2026 · 4 citations
- ClinTutor-R1: Advancing Scalable and Robust One-to-Many Alignment in Clinical Socratic EducationZhitao He, Haolin Yang, Zeyu Qin, Yi FungICML 2026 · 4 citations
- Data-Centric Lessons To Improve Speech-Language PretrainingVishaal Udandarao, Zhiyun Lu, Xuankai Chang, Yongqiang Wang et al.ICLR 2026 · 3 citations
Builds on16
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou et al.ICLR 2024 · 1,197 citations
Related papers
- A Benchmark for Deep Information SynthesisDebjit Paul, Daniel Murphy, Milan Gritta, Ronald Cardenas et al.ICLR 2026 · 1 citation
- USTBench: Benchmarking and Dissecting Spatiotemporal Reasoning Capabilities of LLMs as Urban AgentsSiqi Lai, Yansong Ning, Zirui Yuan, Zhixi Chen et al.ICLR 2026 · 7 citations
- DeepPlanning: Benchmarking Long-Horizon Agentic Planning with Verifiable ConstraintsYinger Zhang, Shutong Jiang, Renhao Li, Jianhong Tu et al.ACL 2026 · 21 citations
- DeepResearch Bench: A Comprehensive Benchmark for Deep Research AgentsMingxuan Du, Benfeng Xu, Chiwei Zhu, Licheng Zhang et al.ICLR 2026 · 250 citations
- LiveNewsBench: Evaluating Web Search Agents with Freshly Curated NewsYunfan Zhang, Kathleen McKeown, Smaranda MuresanICML 2026 · 2 citations
