DIVE: Scaling Diversity in Agentic Task Synthesis for Generalizable Tool Use
Aili Chen, Chi Zhang, Junteng Liu, Jiangjie Chen, Chengyu Du, Yunji Li, Ming Zhong, Qin Wang, Zhengmao Zhu, Jiayuan Song, Ke Ji, Junxian He
摘要
Recent work synthesizes agentic tasks for posttraining tool-using LLMs, yet robust generalization under shifts in tasks and toolsets remains an open challenge. We trace this brittleness to insufficient diversity in synthesized tasks. Scaling diversity is difficult because training requires tasks to remain executable and verifiable, while generalization demands coverage of diverse tool types, toolset combinations, and heterogeneous tool-use patterns. We propose DIVE, an evidencedriven recipe that inverts synthesis order, executing diverse, real-world tools first and reversederiving tasks strictly entailed by the resulting traces, thereby providing grounding by construction. DIVE scales structural diversity along two controllable axes, tool-pool coverage and per-task toolset variety, and an Evidence Collection-Task Derivation loop further induces rich multi-step tool-use patterns across 373 tools in five domains. Training Qwen3-8B on DIVE data (48k SFT + 3.2k RL) improves by +22 average points across 9 OOD benchmarks and outperforms the strongest 8B baseline by +68%. Remarkably, controlled scaling analysis reveals that diversity scaling consistently outperforms quantity scaling for OOD generalization, even with 4× less data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun 等ICLR 2024 · 被引用 716 次
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu 等ACL 2023 · 被引用 540 次
- TravelPlanner: A Benchmark for Real-World Planning with Language AgentsJian Xie, Kai Zhang, Jiangjie Chen, Tinghui Zhu 等ICML 2024 · 被引用 376 次
- MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP ServersZhenting Wang, Qi Chang, Hemani Patel, Shashank Biju 等ICLR 2026 · 被引用 109 次
- The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task ExecutionJunlong Li, Wenshuo Zhao, Jian Zhao, Weihao Zeng 等ICLR 2026 · 被引用 73 次
相关 Paper
- Generalizable End-to-End Tool-Use RL with Synthetic CodeGymWeihua Du, Hailei Gong, Zhan Ling, Kang Liu 等ICLR 2026 · 被引用 13 次
- ScaleEnv: Scaling Environment Synthesis from Scratch for Generalist Interactive Tool-Use Agent TrainingDunwei Tu, Hongyan Hao, Hansi Yang, Yihao Chen 等ICML 2026
- Prismatic Synthesis: Gradient-based Data Diversification Boosts Generalization in LLM ReasoningJaehun Jung, Seungju Han, Ximing Lu, Skyler Hallinan 等NeurIPS 2025 · 被引用 50 次
- Scaling Agentic Capabilities via Grounded Interaction SynthesisWenhang Shi, Jinhao Dong, Yiren Chen, Zhe Zhao 等KDD 2026 · 被引用 1 次
- Less is Enough: Synthesizing Diverse Data in Feature Space of LLMsZhongzhi Li, Xuansheng Wu, Yijiang Li, Lijie Hu 等ICML 2026 · 被引用 1 次
