Benchmarking World-Model Learning with Environment-Level Queries
Archana Warrier, Dat Nguyen, Michelangelo Naim, Moksh Jain, Yichao Liang, Karen Schroeder, Cambridge Yang, Josh Tenenbaum, Sebastian Vollmer, Kevin Ellis, Zenna Tavares
摘要
World models are central to building AI agents capable of flexible reasoning and planning. Yet current evaluations (i) test only properties measurable from observed interactions, such as nextframe prediction or task return, and (ii) do not test whether a learned model supports diverse queries about the environment. In contrast, humans build general-purpose models that can answer many different questions about an environment-including questions that require understanding global structure and counterfactual consequences. We propose WorldTest: a protocol for evaluating whether agents learn models that support multiple environment-level queries-questions whose answers depend on properties of the full environment, not just observed trajectories. Individually, these queries can target properties (e.g., reachability or the effects of interventions) that no single rollout distribution determines. Collectively, they assess model generality across query types. We instantiate WorldTest as AutumnBench, a benchmark of 43 interactive grid-world environments and 129 tasks across three query families for both humans and learning agents. Experiments with 517 human participants and five frontier models show that humans substantially outperform these models, a gap we attribute to differences in exploration and belief updating. AutumnBench provides a framework for evaluating world-model learning in grid-world environments with environmentlevel queries, and WorldTest provides a template for extending such evaluations to richer domains. Benchmarking World-Model Learning with Environment-Level Queries and reports the earliest timestep at which observations diverge from those predicted under the original environment; and (3) planning where the agent produces an action sequence to reach a target configuration, requiring reasoning about long-term consequences of actions. These three task types capture the capabilities illustrated in the kitchen example above: predicting what's in the covered pot, noticing when drawers are reorganized, and planning recipe steps. Together, these task-types yield 129 tasks across all environments covering core world-modeling capabilities including prediction, counterfactual reasoning, and planning. While not exhaustive, we designed AutumnBench to be extensible: new environments and challenges can introduce alternative dynamics (e.g., non-standard physics) or evaluate additional skills such as tool use or analogical reasoning. Finally, we validate AutumnBench through an empirical study involving 517 humans and five state-of-the-art AI models, demonstrating that the benchmark reliably exposes gaps between AI and human world-model learning. Contributions We make the following contributions: • We propose WorldTest, the first theoretical framework for testing world-model learning via environment-level queries through an interaction-based evaluation protocol. • We release AutumnBench, a concrete instantiation of WorldTest in a grid-world setting that operationalizes environment-level queries as automated behavioral tests. Additionally, AutumnBench follows the desiderata for novel games outlined by Ying et al. ( 2025 ) and is designed for easy extensibility. • We evaluate human and AI performance on Autumn-Bench, under the WorldTest framework, demonstrating that AutumnBench reliably exposes gaps between human reasoning and current state-of-the-art models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- Leveraging Procedural Generation to Benchmark Reinforcement LearningKarl Cobbe, Christopher Hesse, Jacob Hilton, John SchulmanICML 2020 · 被引用 685 次
- CLEVRER: Collision Events for Video Representation and ReasoningKexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli 等ICLR 2020 · 被引用 584 次
- In-context Reinforcement Learning with Algorithm DistillationMichael Laskin, Luyu Wang, Junhyuk Oh, Emilio Parisotto 等ICLR 2023 · 被引用 10 次
- Context and Diversity Matter: The Emergence of In-Context Learning in World ModelsFan Wang, ZHIYUAN CHEN, YUXUAN ZHONG, Sunjian Zheng 等ICLR 2026 · 被引用 5 次
相关 Paper
- iWorld-Bench: A Benchmark for Interactive World Models with a Unified Action Generation FrameworkJianjie Fang, Yingshan Lei, Qin Wan, Ziyou Wang 等ICML 2026 · 被引用 9 次
- Implicit Intelligence - Evaluating Agents on What Users Don’t SayVed Sirdeshmukh, Marc WetterICML 2026
- HIS-GPT: Towards 3D Human-In-Scene Multimodal UnderstandingJiahe Zhao, Ruibing Hou, Zejie Tian, Hong Chang 等ICCV 2025 · 被引用 6 次
- ScienceWorld: Is your Agent Smarter than a 5th Grader?Ruoyao Wang, Peter A. Jansen, Marc-Alexandre Côté, Prithviraj AmmanabroluEMNLP 2022 · 被引用 1 次
- ACPBench: Reasoning About Action, Change, and PlanningHarsha Kokel, Michael Katz, Kavitha Srinivas, Shirin SohrabiAAAI 2025 · 被引用 35 次
