AgentSynth: Scalable Task Generation for Generalist Computer-Use Agents
Jingxu Xie, Dylan Xu, Xuandong Zhao, Dawn Song
摘要
We introduce AgentSynth, a scalable and cost-efficient pipeline for automatically synthesizing high-quality tasks and trajectory datasets for generalist computer-use agents. Leveraging information asymmetry, AgentSynth constructs subtasks that are simple during generation but significantly more challenging when composed into long-horizon tasks, enabling the creation of over 6,000 diverse and realistic tasks. A key strength of AgentSynth is its ability to precisely modulate task complexity by varying the number of subtasks. Empirical evaluations show that state-of-the-art LLM agents suffer a steep performance drop, from 18% success at difficulty level 1 to just 4% at level 6, highlighting the benchmark's difficulty and discriminative power. Moreover, our pipeline achieves a low average cost of $0.60 per trajectory, orders of magnitude cheaper than human annotations. Our code and data are available at https://github.com/sunblaze-ucb/AgentSynth
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Scaling Agent Learning via Experience SynthesisZhaorun Chen, Zhuokai Zhao, Kai Zhang, Bo Liu 等ICLR 2026 · 被引用 37 次
- Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement LearningZhaoyang Wang, Canwen Xu, Boyi Liu, Yite Wang 等ICML 2026 · 被引用 25 次
- Scaling Synthetic Task Generation for Agents via ExplorationRam Ramrakhya, Andrew Szot, Omar Attia, Bogdan Mazoure 等ICLR 2026 · 被引用 15 次
- InfiniteWeb: Scalable Web Environment Synthesis for GUI Agent TrainingZiyun Zhang, Zezhou Wang, Xiaoyi Zhang, Zongyu Guo 等ACL 2026 · 被引用 11 次
- WebGym: Scaling Training Environments for Long-Horizon Visual Web Agents with Realistic TasksHao Bai, Alexey Taymanov, Tong Zhang, Aviral Kumar 等CVPR 2026
它引用的顶会 Paper11
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret 等NeurIPS 2024 · 被引用 2,059 次
- WebShop: Towards Scalable Real-World Web Interaction with Grounded Language AgentsShunyu Yao, Howard Chen, John Yang, Karthik NarasimhanNeurIPS 2022 · 被引用 1,477 次
- WizardLM: Empowering Large Pre-Trained Language Models to Follow Complex InstructionsCan Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng 等ICLR 2024 · 被引用 1,206 次
- Large Language Models Cannot Self-Correct Reasoning YetJie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng 等ICLR 2024 · 被引用 858 次
相关 Paper
- Efficient Agent Training for Computer UseYanheng He, Jiahe Jin, Pengfei LiuICLR 2026 · 被引用 15 次
- TaskCraft: Automated Generation of Agentic TasksDingfeng Shi, Jingyi Cao, Qianben Chen, Weichen Sun 等ICLR 2026 · 被引用 49 次
- WebSynthesis: World Model-Guided Monte Carlo Tree Search for Efficient WebAgent Trajectory SynthesisYifei Gao, Junhong Ye, Yifan Yang, Jiaqi Wang 等ACL 2026
- AgentTrek: Agent Trajectory Synthesis via Guiding Replay with Web TutorialsYiheng Xu, Dunjie Lu, Zhennan Shen, Junli Wang 等ICLR 2025
- GTA: Generating Long-horizon Tasks for Web Agents at ScaleTenghao Huang, Kung-Hsiang Huang, Prafulla Kumar Choubey, Yilun Zhou 等ACL 2026
