WebSynthesis: World Model-Guided Monte Carlo Tree Search for Efficient WebAgent Trajectory Synthesis
Yifei Gao, Junhong Ye, Yifan Yang, Jiaqi Wang, Yi Zhang, Ruichen Zhang, Jitao Sang
摘要
Recent advances in large language models (LLMs) have enabled increasingly capable web agents, yet training such agents still relies on high-quality interaction trajectories that are difficult to obtain at scale. We identify two key challenges: (1) Infrastructure Overhead, where network instability and website access restrictions limit data collection scalability; and (2) Constrained Exploration, where irreversible state transitions preclude tree-based search and thus limit trajectory diversity. To address these challenges, we introduce Web-Synthesis, a framework for scalable trajectory synthesis. WebSynthesis employs an LLMbased World Model to simulate state transitions without network dependencies, and integrates Monte Carlo Tree Search to enable reversible exploration over the simulated state space. Experiments on WebArena, WebVoyager, and Mind2Web-Online demonstrate that agents trained exclusively on synthesized trajectories outperform those trained on real-world data, providing a viable alternative to costly real-world data collection. Our code is now available at Github.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Synapse: Trajectory-as-Exemplar Prompting with Memory for Computer ControlLongtao Zheng, Rundong Wang, Xinrun Wang, Bo AnICLR 2024 · 被引用 132 次
- WebVoyager: Building an End-to-End Web Agent with Large Multimodal ModelsHongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu 等ACL 2024 · 被引用 30 次
- Go-Browse: Training Web Agents with Structured ExplorationApurva Gandhi, Graham NeubigICLR 2026 · 被引用 30 次
- GUI-explorer: Autonomous Exploration and Mining of Transition-aware Knowledge for GUI AgentBin Xie, Rui Shao, Gongwei Chen, Kaiwen Zhou 等ACL 2025 · 被引用 28 次
- Web Agents with World Models: Learning and Leveraging Environment Dynamics in Web NavigationHyungjoo Chae, Namyoung Kim, Kai Tzu-iunn Ong, Minju Gwak 等ICLR 2025 · 被引用 2 次
相关 Paper
- Plan-and-Act: Improving Planning of Agents for Long-Horizon TasksLutfi Eren Erdogan, Nicholas Lee, Sehoon Kim, Suhong Moon 等ICML 2025
- AutoWebWorld: Synthesizing Infinite Verifiable Web Environments via Finite State MachinesYifan WU, Yiran Peng, Yiyu Chen, Jianhao Ruan 等ICML 2026 · 被引用 16 次
- WebWorld: A Large-Scale World Model for Web Agent TrainingZikai Xiao, Jianhong Tu, Chuhang Zou, Yuxin Zuo 等ICML 2026 · 被引用 14 次
- AgentTrek: Agent Trajectory Synthesis via Guiding Replay with Web TutorialsYiheng Xu, Dunjie Lu, Zhennan Shen, Junli Wang 等ICLR 2025
- Learn-by-interact: A Data-Centric Framework For Self-Adaptive Agents in Realistic EnvironmentsHongjin Su, Ruoxi Sun, Jinsung Yoon, Pengcheng Yin 等ICLR 2025
