Can Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool Use
Weiming Wu, Song-Lin Lv, Rui Zhu, Zijian Cheng, Lan-Zhe Guo
Abstract
While Large Language Model (LLM) agents demonstrate proficiency in static benchmarks, their deployment in real-world scenarios is hindered by the dynamic nature of user queries, tool sets, and interaction dynamics.To address this generalization gap, we formalize OpenAgent (Tool-Use Agent in Open-World), a problem setting characterized by distributional shifts across query, action, observation, and domain dimensions.To systematically diagnose its impact, we construct a controlled sandbox environment where we define fine-grained environmental shifts across a four-tier hierarchy, Perception , Interaction , Reasoning , and Internalization , and conduct a comprehensive series of experiments. Our analysis yields a series of key insights, demonstrating that agents trained via both Supervised Fine-Tuning (SFT) and Reinforcement Learning suffer from varying degrees of performance degradation when confronting open environmental shifts.Building on these insights, we propose Perturbation-Augmented Fine-Tuning, a disturbance-based intervention strategy for SFT that lays the foundation for enhancing agent robustness and utility in realistic environments. Our code will be released at: https://github.com/LAMDA-NeSy/OpenAgent .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2f5ab85e-829d-47ea-b710-f840fdea500cCited by top-tier papers1
Ask how each one uses itBuilds on3
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou et al.ICLR 2024 · 1,197 citations
- DeepEyes: Incentivizing "Thinking with Images" via Reinforcement LearningZiwei Zheng, Michael Yang, Jack Hong, Chenxiao Zhao et al.ICLR 2026 · 321 citations
- SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-trainingTianzhe Chu, Yuexiang Zhai, Jihan Yang, Shengbang Tong et al.ICML 2025
Related papers
- Don't Just Fine-tune the Agent, Tune the EnvironmentSiyuan Lu, Zechuan Wang, Hongxuan Zhang, Qintong Wu et al.ICLR 2026 · 13 citations
- Does Reinforcement Fine-Tuning Improve Generalization of LLM Agents? An Empirical StudyZhiheng Xi, Xin Guo, Jiaqi Liu, Jiazheng Zhang et al.ICML 2026 · 3 citations
- ToolOmni: Enabling Open-World Tool Use via Agentic learning with Proactive Retrieval and Grounded ExecutionShouzheng Huang, Meishan Zhang, Baotian Hu, Min ZhangACL 2026
- Enhancing Open-Domain Task-Solving Capability of LLMs via Autonomous Tool Integration from GitHubBohan Lyu, Xin Cong, Heyang Yu, Pan Yang et al.ACL 2025
- AgentRefine: Enhancing Agent Generalization through Refinement TuningDayuan Fu, Keqing He, Yejie Wang, Wentao Hong et al.ICLR 2025
