GTA: Generating Long-horizon Tasks for Web Agents at Scale
Tenghao Huang, Kung-Hsiang Huang, Prafulla Kumar Choubey, Yilun Zhou, Muhao Chen, Jonathan May, Chien-Sheng Wu
Abstract
Web agents, which couple language models with browsing and tool-use capabilities, show promise as open web assistants. Yet progress is increasingly limited by the lack of scalable, process-level supervision. Existing benchmarks are largely manually constructed, providing only coarse start-goal annotations without intermediate trajectories, while recent automatic generation efforts remain expensive, biased, and shallow. These limitations prevent reliable training and evaluation of agents that must generalize to realistic, multi-hop, crosspage tasks. We introduce a scalable framework GTA that integrates crawling, retrievalbased seeding, in-context generation, and automated quality control to produce realistic tasks paired with executable trajectories. This design decouples crawling from generation for greater efficiency, grounds tasks in the site graph to enforce compositionality, and ensures dense supervision through deterministic replays and systematic validation. We instantiate the pipeline on over 50 websites covering e-commerce, government, forums, and news, with multilingual and multi-hop coverage. The resulting benchmark reveals a significant human-agent performance gap and enables detailed diagnostics. Our contributions are threefold: (i) formalizing multi-hop web-agent task generation, (ii) proposing an efficient and validated pipeline for automatic data creation, and (iii) releasing a self-evolving benchmark ecosystem where end users can generate upto-date tasks grounded in live web content 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 38cfadda-8a1e-4e71-a101-42079b448641Builds on5
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou et al.ICLR 2024 · 1,197 citations
- AgentOccam: A Simple Yet Strong Baseline for LLM-Based Web AgentsKe Yang, Yao Liu, Sapana Chaudhary, Rasool Fakoor et al.ICLR 2025 · 3 citations
- WebDS: An End-to-End Benchmark for Web-based Data ScienceEthan Hsu, Hong Meng Yam, Ines Bouissou, Aaron Murali John et al.ICLR 2026 · 1 citation
- WebWalker: Benchmarking LLMs in Web TraversalJialong Wu, Wenbiao Yin, Yong Jiang, Zhenglin Wang et al.ACL 2025
- AgentTrek: Agent Trajectory Synthesis via Guiding Replay with Web TutorialsYiheng Xu, Dunjie Lu, Zhennan Shen, Junli Wang et al.ICLR 2025
Related papers
- AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?Ori Yoran, Samuel Joseph Amouyal, Chaitanya Malaviya, Ben Bogin et al.EMNLP 2024 · 5 citations
- WebChain: A Large-Scale Human-Annotated Dataset of Real-World Web Interaction TracesSicheng Fan, Rui Wan, Yifei Leng, Gaoning Liang et al.CVPR 2026 · 4 citations
- Go-Browse: Training Web Agents with Structured ExplorationApurva Gandhi, Graham NeubigICLR 2026 · 30 citations
- WebSynthesis: World Model-Guided Monte Carlo Tree Search for Efficient WebAgent Trajectory SynthesisYifei Gao, Junhong Ye, Yifan Yang, Jiaqi Wang et al.ACL 2026
- WebWorld: A Large-Scale World Model for Web Agent TrainingZikai Xiao, Jianhong Tu, Chuhang Zou, Yuxin Zuo et al.ICML 2026 · 14 citations
