MedAgentGym: A Scalable Agentic Training Environment for Code-Centric Reasoning in Biomedical Data Science
Ran Xu, Yuchen Zhuang, Yishan Zhong, Yue Yu, Zifeng Wang, Xiangru Tang, Hang Wu, May Dongmei Wang, Peifeng Ruan, Donghan Yang, Tao Wang, Guanghua Xiao
Abstract
We introduce MedAgentGym, a scalable and interactive training environment designed to enhance coding-based biomedical reasoning capabilities in large language model (LLM) agents. MedAgentGym comprises 72, 413 task instances across 129 categories derived from 12 authentic real-world biomedical scenarios. Tasks are encapsulated within executable sandbox environments, each featuring detailed task specifications, interactive feedback mechanisms, verifiable ground truth annotations, and scalable training trajectory generation. Extensive benchmarking of 29 LLMs reveals substantial performance disparities in biomedical data science between commercial and open-source LLMs. Leveraging efficient multi-threaded and multi-turn trajectory sampling in MedAgentGym, Med-Copilot achieves performance gains of +43.02% and +45.28% from offline and online reinforcement learning, respectively, demonstrating MedAgentGym as an effective training ground while establishing itself as a cost-effective, privacy-preserving alternative competitive with proprietary LLMs (gpt-4o). By offering a unified execution environment with a comprehensive benchmark and accessible, extensible training resources, MedAgentGym delivers an integrated platform to develop LLM-based coding assistants for advanced biomedical data science.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 23d7034a-3d6c-4f3a-addd-b041ebcc7800Cited by top-tier papers2
- Scaling Agentic Reinforcement Learning for Tool-Integrated Reasoning in VLMsMeng Lu, Ran Xu, Yi Fang, Wenxuan Zhang et al.CVPR 2026 · 15 citations
- SciAgentGym: Benchmarking Multi-Step Scientific Tool-Use in LLM AgentsYujiong Shen, Yajie Yang, Zhiheng Xi, Binze Hu et al.ICML 2026 · 6 citations
Builds on17
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret et al.NeurIPS 2024 · 2,059 citations
- Executable Code Actions Elicit Better LLM AgentsXingyao Wang, Yangyi Chen, Lifan Yuan, Yizhe Zhang et al.ICML 2024 · 436 citations
- MDAgents: An Adaptive Collaboration of LLMs for Medical Decision-MakingYubin Kim, Chanwoo Park, Hyewon Jeong, Yik Siu Chan et al.NeurIPS 2024 · 291 citations
- MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical ReasoningShuyue Stella Li, Vidhisha Balachandran, Shangbin Feng, Jonathan Ilgen et al.NeurIPS 2024 · 215 citations
- MLAgentBench: Evaluating Language Agents on Machine Learning ExperimentationQian Huang, Jian Vora, Percy Liang, Jure LeskovecICML 2024 · 209 citations
Related papers
- AgentGym: Evaluating and Training Large Language Model-based Agents across Diverse EnvironmentsZhiheng Xi, Yiwen Ding, Wenxiang Chen, Boyang Hong et al.ACL 2025 · 20 citations
- AgentBench: Evaluating LLMs as AgentsXiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu et al.ICLR 2024 · 748 citations
- Generalizable End-to-End Tool-Use RL with Synthetic CodeGymWeihua Du, Hailei Gong, Zhan Ling, Kang Liu et al.ICLR 2026 · 13 citations
- InnovatorBench: Evaluating Agents' Ability to Conduct Innovative AI ResearchYunze Wu, Dayuan Fu, Weiye Si, Zhen Huang et al.ICLR 2026 · 9 citations
- WebGym: Scaling Training Environments for Long-Horizon Visual Web Agents with Realistic TasksHao Bai, Alexey Taymanov, Tong Zhang, Aviral Kumar et al.CVPR 2026
