RL-GPT: Integrating Reinforcement Learning and Code-as-policy
Shaoteng Liu, Haoqi Yuan, Minda Hu, Yanwei Li, Yukang Chen, Shu Liu, Zongqing Lu, Jiaya Jia
摘要
Large Language Models (LLMs) have demonstrated proficiency in utilizing various tools by coding, yet they face limitations in handling intricate logic and precise control. In embodied tasks, high-level planning is amenable to direct coding, while low-level actions often necessitate task-specific refinement, such as Reinforcement Learning (RL). To seamlessly integrate both modalities, we introduce a two-level hierarchical framework, RL-GPT, comprising a slow agent and a fast agent. The slow agent analyzes actions suitable for coding, while the fast agent executes coding tasks. This decomposition effectively focuses each agent on specific tasks, proving highly efficient within our pipeline. Our approach outperforms traditional RL methods and existing GPT agents, demonstrating superior efficiency. In the Minecraft game, it rapidly obtains diamonds within a single day on an RTX3090. Additionally, it achieves SOTA performance across all designated MineDojo tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Pre-Trained Multi-Goal Transformers with Prompt Optimization for Efficient Online AdaptationHaoqi Yuan, Yuhui Fu, Feiyang Xie, Zongqing LuNeurIPS 2024 · 被引用 5 次
- Experience Transfer for Multimodal LLM Agents in Minecraft GameChenghao Li, Jun Liu, Songbo Zhang, Huadong Jian 等CVPR 2026 · 被引用 4 次
- Prompted Policy Search: Reinforcement Learning through Linguistic and Numerical Reasoning in LLMsYifan Zhou, Sachin Grover, Mohamed El Mistiri, Kamalesh Kalirathinam 等NeurIPS 2025 · 被引用 3 次
- ReCAPA: Hierarchical Predictive Correction to Mitigate Cascading FailuresXiyin Zeng, Yuyu Sun, Haoyang Li, Shouqiang Liu 等ICLR 2026 · 被引用 2 次
- Cultivating Gaming Sense for Yourself: Making VLMs Gaming ExpertsWenxuan Lu, Jiangyang He, Zhanqiu Zhang, Steven Y. Guo 等ACL 2025 · 被引用 2 次
它引用的顶会 Paper23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch 等ICML 2023 · 被引用 2,601 次
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris 等UIST 2023 · 被引用 1,882 次
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied AgentsWenlong Huang, Pieter Abbeel, Deepak Pathak, Igor MordatchICML 2022 · 被引用 1,539 次
- Large Language Models as OptimizersChengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu 等ICLR 2024 · 被引用 817 次
相关 Paper
- LARM: Large Auto-Regressive Model for Long-Horizon Embodied IntelligenceZhuoling Li, Xiaogang Xu, Zhenhua Xu, Ser-Nam Lim 等ICML 2025
- Bootstrapping Cognitive Agents with a Large Language ModelFeiyu Zhu, Reid G. SimmonsAAAI 2024 · 被引用 14 次
- From Multimodal LLMs to Generalist Embodied Agents: Methods and LessonsAndrew Szot, Bogdan Mazoure, Omar Attia, Aleksei Timofeev 等CVPR 2025
- Describe, Explain, Plan and Select: Interactive Planning with LLMs Enables Open-World Multi-Task AgentsZihao Wang, Shaofei Cai, Guanzhou Chen, Anji Liu 等NeurIPS 2023 · 被引用 178 次
- RF-Agent: Automated Reward Function Design via Language Agent Tree SearchNing Gao, Xiuhui Zhang, Xingyu Jiang, Mukang You 等NeurIPS 2025 · 被引用 8 次
