Scaling Agent Learning via Experience Synthesis
Zhaorun Chen, Zhuokai Zhao, Kai Zhang, Bo Liu, Qi Qi, Yifan Wu, Tarun Kalluri, Xuefei Cao, Yuanhao Xiong, Haibo Tong, Huaxiu Yao, Hengduo Li
Abstract
While reinforcement learning (RL) can empower autonomous agents by enabling self-improvement through interaction, its practical adoption remains challenging due to costly rollouts, limited task diversity, unreliable reward signals, and infrastructure complexity, all of which obstruct the collection of scalable experience data. To address these challenges, we introduce DreamGym, the first unified framework designed to synthesize diverse experiences with scalability in mind to enable effective online RL training for autonomous agents. Rather than relying on expensive real-environment rollouts, DreamGym distills environment dynamics into a reasoning-based experience model that derives consistent state transitions and feedback signals through step-by-step reasoning, enabling scalable agent rollout collection for RL. To improve the stability and quality of transitions, DreamGym leverages an experience replay buffer initialized with offline real-world data and continuously enriched with fresh interactions to actively support agent training. To improve knowledge acquisition, DreamGym adaptively generates new tasks that challenge the current agent policy, enabling more effective online curriculum learning. Experiments across diverse environments and agent backbones demonstrate that DreamGym substantially improves RL training, both in fully synthetic settings and in sim-to-real transfer scenarios. On non-RL-ready tasks like WebArena, DreamGym outperforms all baselines by over 30%. And in RL-ready but costly settings, it matches GRPO and PPO performance using only synthetic interactions. When transferring a policy trained purely on synthetic experiences to real-environment RL, DreamGym yields significant additional performance gains while requiring far fewer real-world interactions, providing a scalable warm-start strategy for general-purpose RL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9ec4fbf2-647d-4a23-b7ae-c9a51273f317Cited by top-tier papers6
- From Word to World: Can Large Language Models be Implicit Text-based World Models?Yixia Li, Hongru Wang, Jiahao Qiu, Zhenfei Yin et al.ACL 2026 · 27 citations
- Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement LearningZhaoyang Wang, Canwen Xu, Boyi Liu, Yite Wang et al.ICML 2026 · 25 citations
- RLAnything: Forge Environment, Policy, and Reward Model in Completely Dynamic RL SystemYinjie Wang, Tianbao Xie, Ke Shen, Mengdi Wang et al.ICML 2026 · 15 citations
- WebWorld: A Large-Scale World Model for Web Agent TrainingZikai Xiao, Jianhong Tu, Chuhang Zou, Yuxin Zuo et al.ICML 2026 · 14 citations
- Can We Predict Before Executing Machine Learning Agents?Jingsheng Zheng, Jintian Zhang, Yujie Luo, Yuren Mao et al.ACL 2026 · 6 citations
Builds on19
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- WebShop: Towards Scalable Real-World Web Interaction with Grounded Language AgentsShunyu Yao, Howard Chen, John Yang, Karthik NarasimhanNeurIPS 2022 · 1,477 citations
Related papers
- WebGym: Scaling Training Environments for Long-Horizon Visual Web Agents with Realistic TasksHao Bai, Alexey Taymanov, Tong Zhang, Aviral Kumar et al.CVPR 2026
- SearchGym: Bootstrapping Real-World Search Agents via Cost-Effective and High-Fidelity Environment SimulationXichen Zhang, Ziyi He, Yinghao Zhu, Sitong Wu et al.ACL 2026 · 2 citations
- RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable EnvironmentsZhiyuan Zeng, Hamish Ivison, Yiping Wang, Lifan Yuan et al.ICML 2026 · 28 citations
- h1: Bootstrapping LLMs to Reason over Longer Horizons via Reinforcement LearningAlesia Ivanova, Sumeet Motwani, Jack Cai, Phil Torr et al.ICML 2026 · 11 citations
- Thinking vs. Doing: Improving Agent Reasoning by Scaling Test-Time InteractionJunhong Shen, Hao Bai, Lunjun Zhang, Yifei Zhou et al.NeurIPS 2025 · 34 citations
