Empowering Efficiency and Efficacy in WebAgent via Enabling Info-Rich Seeking
Zhengwei Tao, Haiyang SHEN, Baixuan Li, Wenbiao Yin, Jialong Wu, Kuan Li, Zhongwang Zhang, Huifeng Yin, Rui Ye, Yong Jiang, Pengjun Xie, Fei Huang
摘要
Large Language Model (LLM)-based agents have emerged as a transformative approach for open-ended problem solving, with information seeking (IS) being a core capability that enables autonomous reasoning and decision-making. While prior research has largely focused on improving retrieval depth, we observe that current IS agents often suffer from low search efficiency, which in turn constrains overall performance. A key factor underlying this inefficiency is the sparsity of target entities in training tasks, which limits opportunities for agents to learn and generalize efficient search behaviors. To address these challenges, we propose WebLeaper, a framework for constructing high-coverage IS tasks and generating efficient solution trajectories. We formulate IS as a tree-structured reasoning problem, enabling a substantially larger set of target entities to be embedded within a constrained context. Leveraging curated Wikipedia tables, we propose three variants for synthesizing IS tasks-Basic, Union, and Reverse-Union-to systematically increase both IS efficiency and efficacy. Finally, we curate training trajectories by retaining only those that are simultaneously accurate and efficient, ensuring that the model is optimized for both correctness and search performance. Extensive experiments on both basic and comprehensive settings, conducted on five IS benchmarks-BrowserComp, GAIA, xbench-DeepSearch, WideSearch, and Seal-0-demonstrate that our method consistently achieves improvements in both effectiveness and efficiency over strong baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- SmartSearch: Process Reward-Guided Query Refinement for Search AgentsTongyu Wen, Guanting Dong, Zhicheng DouSIGIR 2026 · 被引用 13 次
- Nested Browser-Use Learning for Agentic Information SeekingBaixuan Li, Jialong Wu, Wenbiao Yin, Kuan Li 等ACL 2026 · 被引用 6 次
- ET-Agent: Incentivizing Effective Tool-Integrated Reasoning Agent via Behavior CalibrationYifei Chen, Guanting Dong, Zhicheng DouACL 2026 · 被引用 3 次
- Opt-Miner: Empowering Information-Seeking Agent with Tree-Guided Data Synthesis for Optimization ModelingHaoyang Liu, Yuyang Cai, Jie Wang, Xiongwei Han 等ICML 2026
- What Do Agents Learn from Trajectory-SFT: Semantics or Interfaces?Weizheng Gu, Chengze Li, Zhuohao Yu, Mengyuan Sun 等ICML 2026
它引用的顶会 Paper14
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- A-Mem: Agentic Memory for LLM AgentsWujiang Xu, Zujie Liang, Kai Mei, Hang Gao 等NeurIPS 2025 · 被引用 1,138 次
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun 等ICLR 2024 · 被引用 716 次
- WebThinker: Empowering Large Reasoning Models with Deep Research CapabilityXiaoxi Li, Jiajie Jin, Guanting Dong, Hongjin Qian 等NeurIPS 2025 · 被引用 354 次
- SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated ReasoningZhenghai Xue, Longtao Zheng, Qian Liu, Yingru Li 等ICLR 2026 · 被引用 152 次
相关 Paper
- WebShaper: Agentically Data Synthesizing via Information-Seeking FormalizationZhengwei Tao, Jialong Wu, Wenbiao Yin, Pu Wu 等ICLR 2026 · 被引用 115 次
- Open Data Synthesis for Deep ResearchZiyi Xia, Kun Luo, Hongjin Qian, Siqi Bao 等ICLR 2026 · 被引用 14 次
- Learning to Retrieve from Agent TrajectoriesYuqi Zhou, Sunhao Dai, Changle Qu, Liang Pang 等SIGIR 2026
- WebClipper: Efficient Evolution of Web Agents with Graph-based Trajectory PruningJunjie Wang, Zequn Xie, Dan Yang, Jie Feng 等ACL 2026
- WebWatcher: Breaking New Frontiers of Vision-Language Deep Research AgentXinyu Geng, Peng Xia, Zhen Zhang, Xinyu Wang 等ICLR 2026 · 被引用 79 次
