Current Agents Fail to Leverage World Model as Tool for Foresight
Cheng Qian, Emre Can Acikgoz, Bingxuan Li, Xiusi Chen, Yuji Zhang, Bingxiang He, Qinyu Luo, Gokhan Tur, Dilek Hakkani-Tür, Yunzhu Li, Heng Ji
摘要
Agents built on vision-language models increasingly face tasks that demand anticipating future states rather than relying on short-horizon reasoning. Generative world models offer a promising remedy: agents could use them as external simulators to foresee outcomes before acting. This paper empirically examines whether current agents can leverage such world models as tools to enhance their cognition. Across diverse agentic and visual question answering tasks, we observe that some agents rarely invoke simulation (fewer than 1%), frequently misuse predicted rollouts (approximately 15%), and often exhibit inconsistent or even degraded performance (up to 5%) when simulation is available or enforced. Attribution analysis further indicates that the primary bottleneck lies in the agents'capacity to decide when to simulate, how to interpret predicted outcomes, and how to integrate foresight into downstream reasoning. These findings underscore the need for mechanisms that foster calibrated, strategic interaction with world models, paving the way toward more reliable anticipatory cognition in future agent systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper7
- Reasoning with Language Model is Planning with World ModelShibo Hao, Yi Gu, Haodi Ma, Joshua Jiahua Hong 等EMNLP 2023 · 被引用 109 次
- Energy-Based Transformers are Scalable Learners and ThinkersAlexi Gladstone, Ganesh Nanduru, Md Mofijul Islam, Peixuan Han 等ICLR 2026 · 被引用 38 次
- Pre-Trained Video Generative Models as World SimulatorsHaoran He, Yang Zhang, Liang Lin, Zhongwen Xu 等AAAI 2026 · 被引用 32 次
- 3DSRBENCH: A Comprehensive 3D Spatial Reasoning BenchmarkWufei Ma, Haoyu Chen, Guofeng Zhang, Yu-Cheng Chou 等ICCV 2025 · 被引用 15 次
- An Empirical Analysis on Spatial Reasoning Capabilities of Large Multimodal ModelsFatemeh Shiri, Xiao-Yu Guo, Mona Far, Xin Yu 等EMNLP 2024 · 被引用 7 次
相关 Paper
- VLAW: Iterative Co-Improvement of Vision-Language-Action Policy and World ModelYanjiang Guo, Tony Lee, Lucy Xiaoyang Shi, Jianyu Chen 等ICML 2026 · 被引用 29 次
- Grounded Answers for Multi-agent Decision-making Problem through Generative World ModelZeyang Liu, Xinrui Yang, Shiguang Sun, Long Qian 等NeurIPS 2024 · 被引用 10 次
- R-WoM: Retrieval-augmented World Model For Computer-use AgentsKai Mei, Jiang Guo, Shuaichen Chang, Mingwen Dong 等ICLR 2026 · 被引用 12 次
- ForeAct: Steering Your VLA with Efficient Visual Foresight PlanningZhuoyang Zhang, Shang Yang, Qinghao Hu, Luke J. Huang 等CVPR 2026 · 被引用 7 次
- From Word to World: Can Large Language Models be Implicit Text-based World Models?Yixia Li, Hongru Wang, Jiahao Qiu, Zhenfei Yin 等ACL 2026 · 被引用 27 次
