How to Train Your LLM Web Agent: A Statistical Diagnosis
Dheeraj Vattikonda, Santhoshi Ravichandran, Emiliano Penaloza, Hadi Nekoei, Thibault Le Sellier de Chezelles, Megh Thakkar, Nicolas Gontier, Miguel Muñoz-Mármol, Sahar Omidi Shayegan, Stefania Raimondo, Steve (Xue) Liu, Alexandre Drouin
Abstract
LLM-based web agents have recently made significant progress, but much of it has occurred in closed-source systems, widening the gap with open-source alternatives. Progress has been held back by two key challenges: first, a narrow focus on single-step tasks that overlooks the complexity of multi-step web interactions; and second, the high compute costs required to post-train LLM-based web agents. To address this, we present the first statistically grounded study on compute allocation for LLM web-agent post-training. Our approach uses a two-stage pipeline, training a Llama 3.1 8B student to imitate a Llama 3.3 70B teacher via supervised fine-tuning (SFT), followed by on-policy reinforcement learning. We find this process highly sensitive to hyperparameter choices, making exhaustive sweeps impractical. To spare others from expensive trial-and-error, we sample 1,370 configurations and use bootstrapping to estimate effective hyperparameters. Our results show that combining SFT with on-policy RL consistently outperforms either approach alone on both WorkArena and MiniWob++. Further, this strategy requires only 55% of the compute to match the peak performance of pure SFT on MiniWob++, effectively pushing the compute-performance Pareto frontier, and is the only strategy that can close the gap with closed-source models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 50053173-409c-4651-aaab-052b20880aedCited by top-tier papers3
- OpenApps: Simulating Environment Variations to Measure UI Agent ReliabilityKaren Ullrich, Jingtong Su, Claudia Shi, Arjun Subramonian et al.ICLR 2026 · 10 citations
- On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon LengthSunghwan Kim, Junhee Cho, Beong-woo Kwak, Taeyoon Kwon et al.ICML 2026 · 3 citations
- Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical ReasoningBowen Ding, Yuhan Chen, Jiayang Lyu, Jiyao Yuan et al.ACL 2026
Builds on8
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou et al.ICLR 2024 · 1,197 citations
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
- WorkArena: How Capable are Web Agents at Solving Common Knowledge Work Tasks?Alexandre Drouin, Maxime Gasse, Massimo Caccia, Issam H. Laradji et al.ICML 2024 · 188 citations
Related papers
- WebAgent-R1: Training Web Agents via End-to-End Multi-Turn Reinforcement LearningZhepei Wei, Wenlin Yao, Yao Liu, Weizhi Zhang et al.EMNLP 2025 · 1 citation
- WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement LearningZehan Qi, Xiao Liu, Iat Long Iong, Hanyu Lai et al.ICLR 2025
- Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language ModelsMichael Noukhovitch, Shengyi Huang, Sophie Xhonneux, Arian Hosseini et al.ICLR 2025
- Retaining by Doing: The Role of On-Policy Data in Mitigating ForgettingHoward Chen, Noam Razin, Karthik Narasimhan, Danqi ChenICML 2026
- IsoCompute Playbook: Optimally Scaling Sampling Compute for LLM RLZhoujun Cheng, Yutao Xie, Yuxiao Qu, Amrith Setlur et al.ICML 2026
