WebSailor-V2: Bridging the Chasm to Proprietary Agents via Synthetic Data and Scalable Reinforcement Learning
Kuan Li, Zhongwang Zhang, Huifeng Yin, Rui Ye, Yida Zhao, Liwen Zhang, Litu Ou, Dingchu Zhang, Xixi Wu, Xinmiao Yu, Jialong Wu, Xinyu Wang
Abstract
To significantly advance the capabilities of open-source web agents, we present WebSailor-V2, a complete post-training pipeline encompassing data construction, Supervised Fine-Tuning (SFT), and Reinforcement Learning (RL). Our methodology features two key innovations: (1) On the data front, we developed SailorFog-QA-2, a novel dataset built from a densely interconnected knowledge graph that introduces a wide variety of uncertainties beyond simple obfuscation, fostering more sophisticated reasoning. ( 2 ) For training, we engineered a dual-environment RL framework, combining a high-fidelity simulator for rapid, low-cost algorithmic iteration with a robust, managed real-world environment for stable final policy training, all integrated within a symbiotic data-policy feedback loop. Trained on the Qwen3-30B-A3B model, WebSailor-V2 achieves state-of-the-art results, scoring 35.3 on BrowseComp-EN, 44.1 on BrowseComp-ZH, and 30.6 on Humanity's Last Exam (HLE). Notably, our 30B-A3B MOE agent significantly outperforms all existing open-source agents and surpasses even the 671B DeepSeek-V3.1, demonstrating performance competitive with leading proprietary systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1629fc41-1bd8-4360-bb25-5811c8b3986bCited by top-tier papers18
- Scaling Long-Horizon Agent via Context FoldingWeiwei Sun, Lu Miao, Zhan Ling, Kang Liu et al.ICML 2026 · 104 citations
- WebWeaver: Structuring Web-Scale Evidence with Dynamic Outlines for Open-Ended Deep ResearchZijian Li, Xin Guan, Bo Zhang, Shen Huang et al.ICLR 2026 · 48 citations
- Scaling Agents via Continual Pre-trainingLiangcai Su, Zhen Zhang, Guangyu Li, Zhuo Chen et al.ICLR 2026 · 46 citations
- AgentFold: Long-Horizon Web Agents with Proactive Context FoldingRui Ye, Zhongwang Zhang, Kuan Li, Huifeng Yin et al.ICLR 2026 · 45 citations
- Search Self-Play: Pushing the Frontier of Agent Capability without SupervisionHongliang Lu, Yuhang Wen, Pengyu Cheng, Ruijin Ding et al.ICLR 2026 · 35 citations
Builds on8
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun et al.ICLR 2024 · 716 citations
- WebThinker: Empowering Large Reasoning Models with Deep Research CapabilityXiaoxi Li, Jiajie Jin, Guanting Dong, Hongjin Qian et al.NeurIPS 2025 · 354 citations
- DeepResearch Bench: A Comprehensive Benchmark for Deep Research AgentsMingxuan Du, Benfeng Xu, Chiwei Zhu, Licheng Zhang et al.ICLR 2026 · 250 citations
- MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon AgentsZijian Zhou, Ao Qu, Zhaoxuan Wu, Sunghwan Kim et al.ICLR 2026 · 223 citations
Related papers
- WebWatcher: Breaking New Frontiers of Vision-Language Deep Research AgentXinyu Geng, Peng Xia, Zhen Zhang, Xinyu Wang et al.ICLR 2026 · 79 citations
- SearchGym: Bootstrapping Real-World Search Agents via Cost-Effective and High-Fidelity Environment SimulationXichen Zhang, Ziyi He, Yinghao Zhu, Sitong Wu et al.ACL 2026 · 2 citations
- WebAggregator: Enhancing Compositional Reasoning Capabilities of Deep Research Agent Foundation ModelsRui Wang, Ce Zhang, Jun-Yu Ma, Jianshu Zhang et al.ACL 2026 · 4 citations
- How to Train Your LLM Web Agent: A Statistical DiagnosisDheeraj Vattikonda, Santhoshi Ravichandran, Emiliano Penaloza, Hadi Nekoei et al.NeurIPS 2025 · 19 citations
- DRIVE: Best Data Scheduling Practices for Reinforcement Learning with Verifiable Reward in Competitive Code GenerationSpeed Zhu, Chuheng Zhang, Jianwei Cai, Guang Chen et al.ICML 2026
