R-WoM: Retrieval-augmented World Model For Computer-use Agents
Kai Mei, Jiang Guo, Shuaichen Chang, Mingwen Dong, Dongkyu Lee, Xing Niu, Jiarong Jiang
摘要
Large Language Models (LLMs) can serve as world models to enhance agent decision-making in digital environments by simulating future states and predicting action outcomes, potentially eliminating costly trial-and-error exploration. However, this capability is fundamentally limited by LLM's tendency to hallucination and their reliance on static training knowledge, which could lead to compounding errors that inhibit long-horizon simulations. To systematically investigate whether LLMs are appropriate for world modeling, we probe two core capabilities of world models -- future state prediction and reward estimation -- through three tasks: next-state identification, full-procedure planning alignment, and milestone transition recognition. Our analysis shows that while LLMs effectively capture immediate next states and identify meaningful state transitions, their performance rapidly degrades in full-procedure planning. This highlights LLMs’ limitations in reliably modeling environment dynamics over long horizons. To address these limitations, we propose the Retrieval-augmented World Model (R-WoM), which grounds LLM simulations by incorporating factual, up-to-date knowledge retrieved from external tutorials. Experiments show that R-WoM achieves relative improvements of up to 23.4% and 16.3% on the subsets of OSWorld and Webarena compared to baselines, with particular advantage in longer-horizon simulations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- OS-Symphony: A Holistic Framework for Robust and Generalist Computer-Using AgentsBowen Yang, Kaiming Jin, Zhenyu Wu, Zhaoyang Liu 等ACL 2026 · 被引用 15 次
- Reasoning over Precedents Alongside Statutes: Case-Augmented Deliberative Alignment for LLM SafetyCan Jin, Rui Wu, Tong Che, Qixin Zhang 等ACL 2026 · 被引用 3 次
它引用的顶会 Paper11
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 被引用 1,170 次
- Generative Judge for Evaluating AlignmentJunlong Li, Shichao Sun, Weizhe Yuan, Run-Ze Fan 等ICLR 2024 · 被引用 173 次
- Reasoning with Language Model is Planning with World ModelShibo Hao, Yi Gu, Haodi Ma, Joshua Jiahua Hong 等EMNLP 2023 · 被引用 109 次
- ComputerRL: Scaling End-to-End Online Reinforcement Learning for Computer Use AgentsHanyu Lai, Xiao Liu, Yanxiao Zhao, Han Xu 等ICLR 2026 · 被引用 45 次
相关 Paper
- Web Agents with World Models: Learning and Leveraging Environment Dynamics in Web NavigationHyungjoo Chae, Namyoung Kim, Kai Tzu-iunn Ong, Minju Gwak 等ICLR 2025 · 被引用 2 次
- Agent Planning with World Knowledge ModelShuofei Qiao, Runnan Fang, Ningyu Zhang, Yuqi Zhu 等NeurIPS 2024 · 被引用 95 次
- From Word to World: Can Large Language Models be Implicit Text-based World Models?Yixia Li, Hongru Wang, Jiahao Qiu, Zhenfei Yin 等ACL 2026 · 被引用 27 次
- WebEvolver: Enhancing Web Agent Self-Improvement with Co-evolving World ModelTianqing Fang, Hongming Zhang, Zhisong Zhang, Kaixin Ma 等EMNLP 2025
- Scaling Autonomous Agents via Automatic Reward Modeling And PlanningZhenfang Chen, Delin Chen, Rui Sun, Wenjun Liu 等ICLR 2025
