Agent Learning via Early Experience
Kai Zhang, Xiangchao Chen, Bo Liu, Tianci Xue, Zeyi Liao, Zhihan Liu, xiyao wang, Yuting Ning, Zhaorun Chen, Xiaohan Fu, Jian Xie, Yuxuan Sun
Abstract
A long-term goal of language agents is to learn and improve through their own experience, ultimately outperforming humans in complex, real-world tasks. However, training agents from experience data with reinforcement learning remains difficult in many environments, which either lack verifiable rewards (e.g., websites) or require inefficient long-horizon rollouts (e.g., multi-turn tool use). As a result, most current agents rely on supervised fine-tuning on expert data, which is challenging to scale and generalizes poorly. This limitation stems from the nature of expert demonstrations: they capture only a narrow range of scenarios, and expose the agent to limited environment diversity. We address this limitation with a middle-ground paradigm we call early experience : interaction data generated by the agent's own actions, where the resulting future states serve as supervision without reward signals. Within this paradigm, we study two strategies of using such data: (1) implicit world modeling, which uses collected states to ground the policy in environment dynamics; and (2) self-reflection, where the agent learns from its suboptimal actions to improve reasoning and decision-making. Evaluation across eight diverse environments and multiple model families shows that our approaches consistently improve effectiveness and out-of-domain generalization, highlighting the value of early experience. Moreover, in environments with verifiable rewards, our results provide promising signals that early experience offers a strong foundation for subsequent reinforcement learning, making it a practical bridge between imitation learning and fully experience-driven agents.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2f060c88-3f15-4a6e-be45-e2cd45fae2baCited by top-tier papers13
- Evolving AgentsLeonardo RanaldiACL 2026 · 227 citations
- On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language ModelsCharlie Zhang, Graham Neubig, Xiang YueICML 2026 · 58 citations
- Expanding the Capabilities of Reinforcement Learning via Text FeedbackYuda Song, Lili Chen, Fahim Tajwar, REMI MUNOS et al.ICML 2026 · 41 citations
- Scaling Agent Learning via Experience SynthesisZhaorun Chen, Zhuokai Zhao, Kai Zhang, Bo Liu et al.ICLR 2026 · 37 citations
- From Word to World: Can Large Language Models be Implicit Text-based World Models?Yixia Li, Hongru Wang, Jiahao Qiu, Zhenfei Yin et al.ACL 2026 · 27 citations
Builds on29
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
Related papers
- Don't Just Fine-tune the Agent, Tune the EnvironmentSiyuan Lu, Zechuan Wang, Hongxuan Zhang, Qintong Wu et al.ICLR 2026 · 13 citations
- Passive learning of active causal strategies in agents and language modelsAndrew K. Lampinen, Stephanie C. Y. Chan, Ishita Dasgupta, Andrew J. Nam et al.NeurIPS 2023 · 30 citations
- The Confidence Dichotomy: Analyzing and Mitigating Miscalibration in Tool-Use AgentsWeihao Xuan, Qingcheng Zeng, Heli Qi, Yunze Xiao et al.ACL 2026 · 4 citations
- When Greedy Wins: Emergent Exploitation Bias in Meta-Bandit LLM TrainingSanxing Chen, Xiaoyin Chen, Yukun Huang, Roy Xie et al.ICLR 2026 · 3 citations
- Efficient Reinforcement Learning by Guiding World Models with Non-Curated DataYi Zhao, Aidan Scannell, Wenshuai Zhao, Yuxin Hou et al.ICLR 2026 · 2 citations
