AgentGym: Evaluating and Training Large Language Model-based Agents across Diverse Environments
Zhiheng Xi, Yiwen Ding, Wenxiang Chen, Boyang Hong, Honglin Guo, Junzhe Wang, Xin Guo, Dingwen Yang, Chenyang Liao, Wei He, Songyang Gao, Lu Chen
Abstract
Large language models (LLMs) have emerged as a promising foundation to build generally-capable agents (LLM-based agents) that can handle multi-turn decision-making tasks across various environments. However, the community lacks a unified interactive framework that covers diverse environments for comprehensive evaluation of agents, and enables exploration and learning for their self-improvement. To address this, we propose A GENT G YM , a framework featuring 7 real-world scenarios, 14 environments, and 89 tasks for unified, real-time, and concurrent agent interaction. We construct expanded instruction set, high-quality trajectories, and comprehensive benchmarking suite for developing LLM-based agents. Moreover, A GENT G YM supports interactive exploration and learning for agents through multi-turn interactions and real-time feedback. Based on A GENT G YM , we take the initial step to develop LLM-based agents that can handle diverse tasks via methods like self-improvement or reinforcement learning. Experimental re-sults show that the trained agents can achieve results comparable to commercial models. We hope our work can help the community develop more advanced LLM-based agents. We release the code, dataset, benchmark, and checkpoints at https://agentgym.github.io/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers14
- Agent Data Protocol: Unifying Datasets for Diverse, Effective Fine-tuning of LLM AgentsYueqi Song, Ketan Ramaneti, Zaid Sheikh, Ziru Chen et al.ICLR 2026 · 18 citations
- Uncertainty Quantification in LLM Agents: Foundations, Emerging Challenges, and OpportunitiesChangdae Oh, Seongheon Park, To Eun Kim, Jiatong Li et al.ACL 2026 · 8 citations
- RoboAgent: Chaining Basic Capabilities for Embodied Task PlanningPeiran Xu, Jiaqi Zheng, Yadong MuCVPR 2026 · 6 citations
- SciAgentGym: Benchmarking Multi-Step Scientific Tool-Use in LLM AgentsYujiong Shen, Yajie Yang, Zhiheng Xi, Binze Hu et al.ICML 2026 · 6 citations
- Does Reinforcement Fine-Tuning Improve Generalization of LLM Agents? An Empirical StudyZhiheng Xi, Xin Guo, Jiaqi Liu, Jiazheng Zhang et al.ICML 2026 · 3 citations
Builds on12
- WebShop: Towards Scalable Real-World Web Interaction with Grounded Language AgentsShunyu Yao, Howard Chen, John Yang, Karthik NarasimhanNeurIPS 2022 · 1,477 citations
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou et al.ICLR 2024 · 1,197 citations
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 1,126 citations
- ALFWorld: Aligning Text and Embodied Environments for Interactive LearningMohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk et al.ICLR 2021 · 819 citations
- LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language ModelsChan Hee Song, Brian M. Sadler, Jiaman Wu, Wei-Lun Chao et al.ICCV 2023 · 685 citations
Related papers
- AgentGym-RL: An Open-Source Framework to Train LLM Agents for Long-Horizon Decision Making via Multi-Turn RLZhiheng Xi, Jixuan Huang, Chenyang Liao, Baodai Huang et al.ICLR 2026
- LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language ModelsMarwa Abdulhai, Isadora White, Charlie Victor Snell, Charles Sun et al.ICML 2025
- GEM: A Gym for Generalist LLMsZichen Liu, Anya Sims, Keyu Duan, Changyu Chen et al.ICLR 2026
- AgentBench: Evaluating LLMs as AgentsXiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu et al.ICLR 2024 · 748 citations
- VirtualEnv: A Platform for Embodied AI ResearchKabir Swain, Sijie Han, Ayush Raina, Jin Zhang et al.AAAI 2026
