Auto MC-Reward: Automated Dense Reward Design with Large Language Models for Minecraft
Hao Li, Xue Yang, Zhaokai Wang, Xizhou Zhu, Jie Zhou, Yu Qiao, Xiaogang Wang, Hongsheng Li, Lewei Lu, Jifeng Dai
摘要
5 SenseTime Research https://yangxue0827.github.io/auto_mc-reward.html Figure 1. Overview of our Auto MC-Reward. Auto MC-Reward consists of three key LLM-based components: Reward Designer, Reward Critic, and Trajectory Analyzer. A suitable dense reward function is iterated through the continuous interaction between the agent and the environment for reinforcement learning training of specific tasks, so that the model can better complete the task. An example of exploring diamond ore is shown in the figure: i) Trajectory Analyzer finds that the agent dies from lava in the failed trajectory, and then gives suggestion for punishment when encountering lava; ii) Reward Designer adopts the suggestion and updates the reward function; iii) The revised reward function passes the review of Reward Critic, and finally the agent avoids the lava by turning left.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- RL-GPT: Integrating Reinforcement Learning and Code-as-policyShaoteng Liu, Haoqi Yuan, Minda Hu, Yanwei Li 等NeurIPS 2024 · 被引用 48 次
- Multi-Objective Evolution of Heuristic Using Large Language ModelShunyu Yao, Fei Liu, Xi Lin, Zhichao Lu 等AAAI 2025 · 被引用 48 次
- A Decision-Language Model (DLM) for Dynamic Restless Multi-Armed Bandit Tasks in Public HealthNikhil Behari, Edwin Zhang, Yunfan Zhao, Aparna Taneja 等NeurIPS 2024 · 被引用 39 次
- How Do Multimodal Large Language Models Handle Complex Multimodal Reasoning? Placing Them in an Extensible Escape GameZiyue Wang, Yurui Dong, Fuwen Luo, Minyuan Ruan 等ICCV 2025 · 被引用 9 次
- Experience Transfer for Multimodal LLM Agents in Minecraft GameChenghao Li, Jun Liu, Songbo Zhang, Huadong Jian 等CVPR 2026 · 被引用 4 次
它引用的顶会 Paper6
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Guiding Pretraining in Reinforcement Learning with Large Language ModelsYuqing Du, Olivia Watkins, Zihan Wang, Cédric Colas 等ICML 2023 · 被引用 257 次
- Describe, Explain, Plan and Select: Interactive Planning with LLMs Enables Open-World Multi-Task AgentsZihao Wang, Shaofei Cai, Guanzhou Chen, Anji Liu 等NeurIPS 2023 · 被引用 178 次
- TorchRL: A data-driven decision-making library for PyTorchAlbert Bou, Matteo Bettini, Sebastian Dittert, Vikash Kumar 等ICLR 2024 · 被引用 77 次
- EAGER: Asking and Answering Questions for Automatic Reward Shaping in Language-guided RLThomas Carta, Pierre-Yves Oudeyer, Olivier Sigaud, Sylvain LamprierNeurIPS 2022 · 被引用 35 次
相关 Paper
- Text2Reward: Reward Shaping with Language Models for Reinforcement LearningTianbao Xie, Siheng Zhao, Chen Henry Wu, Yitao Liu 等ICLR 2024 · 被引用 142 次
- Empowering Large Language Model Agent through Step-Level Self-Critique and Self-TrainingYuanzhao Zhai, Huanxi Liu, Zhuo Zhang, Tong Lin 等SIGIR 2025 · 被引用 2 次
- Code as Reward: Empowering Reinforcement Learning with VLMsDavid Venuto, Mohammad Sami Nur Islam, Martin Klissarov, Doina Precup 等ICML 2024 · 被引用 29 次
- ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-TrainingYu Liang, Liangxin Liu, Longzheng Wang, Yan Wang 等ACL 2026
- ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree SearchDan Zhang, Sining Zhoubian, Ziniu Hu, Yisong Yue 等NeurIPS 2024 · 被引用 527 次
