Learning to Adapt: Self-Improving Web Agent via Cognitive-Aware Exploration
Weile Chen, Bingchen Miao, Qifan Yu, Wendong Bu, Guoming Wang, Wenqiao Zhang, Shengyu Zhang, Juncheng Li, Siliang Tang
Abstract
Recent advances in Multimodal Large Language Models (MLLMs) have led to promising progress in web agents. However, existing web agents often rely on handcrafted execution pipelines or expensive expert trajectories, limiting their adaptability to complex, dynamic environments. To address these challenges, we propose SCALE (Self-Cognitive-Aware Learning and Exploration), which leverages three adversarial roles, Selector, Predictor, and Judger to autonomously discover the agent's limitations and expand its cognitive boundaries through environmental exploration. Moreover, we propose SCALE-Hop, a graph exploration strategy that facilitates global planning and helps agents avoid local exploration traps. To further support learning, we construct SCALE-20k, a large-scale dataset collected from 19 real-world websites, containing diverse task types and structured demonstrations generated from SCALE's exploration traces. Experimental results show that our approach significantly improves the performance and generalization of multiple MLLMs in various web environments. Our framework offers a scalable and generalizable solution for building truly autonomous and adaptive web agents.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 204e5694-554f-4bc0-b0df-5b28a2a8d1f6Builds on19
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou et al.ICLR 2024 · 1,197 citations
- GPT-4V(ision) is a Generalist Web Agent, if GroundedBoyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun et al.ICML 2024 · 496 citations
- OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task SynthesisQiushi Sun, Kanzhi Cheng, Zichen Ding, Chuanyang Jin et al.ACL 2025 · 114 citations
- Grounded Reinforcement Learning for Visual ReasoningGabriel Sarch, Snigdha Saha, Naitik Khandelwal, Ayush Jain et al.NeurIPS 2025 · 90 citations
Related papers
- AdaptAgent: Adapting Multimodal Web Agents with Few-Shot Learning from Human DemonstrationsGaurav Verma, Rachneet Kaur, Nishan Srishankar, Zhen Zeng et al.ACL 2025 · 19 citations
- Empowering GUI Agents via Autonomous Experience Exploration and Hindsight Experience Utilization for Task PlanningTianyi Men, Zhuoran Jin, Pengfei Cao, Yubo Chen et al.ACL 2026
- Go-Browse: Training Web Agents with Structured ExplorationApurva Gandhi, Graham NeubigICLR 2026 · 30 citations
- On the Multi-turn Instruction Following for Conversational Web AgentsYang Deng, Xuan Zhang, Wenxuan Zhang, Yifei Yuan et al.ACL 2024 · 5 citations
- Graph2Eval: Automatic Multimodal Task Generation for Agents via Knowledge GraphsYurun Chen, Xueyu Hu, Yuhan Liu, Ziqi Wang et al.CVPR 2026 · 7 citations
