AgentTrek: Agent Trajectory Synthesis via Guiding Replay with Web Tutorials
Yiheng Xu, Dunjie Lu, Zhennan Shen, Junli Wang, Zekun Wang, Yuchen Mao, Caiming Xiong, Tao Yu
摘要
Graphical User Interface (GUI) agents can automate complex tasks across digital environments, but their development is hindered by the scarcity of high-quality trajectory data for training. Existing approaches rely on expensive human annotation, making them unsustainable at scale. We propose AgentTrek, a scalable data synthesis pipeline that generates web agent trajectories by leveraging publicly available tutorials. Our three-stage method: (1) automatically harvests and filters tutorial-like texts from the internet using a specialized classification model, (2) transforms these texts into structured task specifications with step-bystep instructions, and (3) employs a visual-language model (VLM) agent to execute these instructions in real environments, while a VLM-based evaluator verifies trajectory correctness. The synthesized trajectories encompass multiple modalities, including text-based HTML observations with function-calling API actions, and vision-based screenshot observations with pixel-level actions. This multimodal data, enriched with chain-of-thought reasoning, enables agents to achieve state-of-the-art performance on both textual web browsing benchmarks (e.g., We-bArena) and visual web grounding and browsing benchmarks (e.g., ScreenSpot Web and Multimodal Mind2Web). Furthermore, our fully automated approach significantly reduces data collection costs, achieving a cost of just $0.55 per highquality trajectory without human annotators. Our work demonstrates that guided replay using web tutorials is a practical and scalable strategy for training advanced GUI agents, paving the way for more capable and autonomous digital assistants.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper32
- OpenCUA: Open Foundations for Computer-Use AgentsXinyuan Wang, Bowen Wang, Dunjie Lu, Junlin Yang 等NeurIPS 2025 · 被引用 151 次
- SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from ExperienceZEYI SUN, Ziyu Liu, Yuhang Zang, Yuhang Cao 等ICML 2026 · 被引用 58 次
- Earth-Agent: Unlocking the Full Landscape of Earth Observation with AgentsPeilin Feng, Zhutao Lv, Junyan Ye, Xiaolei Wang 等ICLR 2026 · 被引用 49 次
- Scaling Agent Learning via Experience SynthesisZhaorun Chen, Zhuokai Zhao, Kai Zhang, Bo Liu 等ICLR 2026 · 被引用 37 次
- AgentSynth: Scalable Task Generation for Generalist Computer-Use AgentsJingxu Xie, Dylan Xu, Xuandong Zhao, Dawn SongICLR 2026 · 被引用 36 次
它引用的顶会 Paper6
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- WebShop: Towards Scalable Real-World Web Interaction with Grounded Language AgentsShunyu Yao, Howard Chen, John Yang, Karthik NarasimhanNeurIPS 2022 · 被引用 1,477 次
- GPT-4V(ision) is a Generalist Web Agent, if GroundedBoyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun 等ICML 2024 · 被引用 496 次
- Patch n' Pack: NaViT, a Vision Transformer for any Aspect Ratio and ResolutionMostafa Dehghani, Basil Mustafa, Josip Djolonga, Jonathan Heek 等NeurIPS 2023 · 被引用 303 次
- ZeroGen: Efficient Zero-shot Learning via Dataset GenerationJiacheng Ye, Jiahui Gao, Qintong Li, Hang Xu 等EMNLP 2022 · 被引用 96 次
相关 Paper
- Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent PretrainingWeimin Xiong, Shuhao Gu, Bowen Ye, Zihao Yue 等ICML 2026 · 被引用 2 次
- VideoAgentTrek: Computer-Use Pretraining from Unlabeled VideosDunjie Lu, Yiheng Xu, Junli Wang, Haoyuan Wu 等ICLR 2026 · 被引用 20 次
- TongUI: Internet-Scale Trajectories from Multimodal Web Tutorials for Generalized GUI AgentsBofei Zhang, Zirui Shang, Zhi Gao, Wang Zhang 等AAAI 2026 · 被引用 26 次
- WebChain: A Large-Scale Human-Annotated Dataset of Real-World Web Interaction TracesSicheng Fan, Rui Wan, Yifei Leng, Gaoning Liang 等CVPR 2026 · 被引用 4 次
- WebSTAR: Scalable Data Synthesis for Computer Use Agents with Step-Level FilteringYifei He, Pranit Chawla, Yaser Souri, Subhojit Som 等ACL 2026 · 被引用 5 次
