LARM: Large Auto-Regressive Model for Long-Horizon Embodied Intelligence
Zhuoling Li, Xiaogang Xu, Zhenhua Xu, Ser-Nam Lim, Hengshuang Zhao
摘要
Recent embodied agents are primarily built based on reinforcement learning (RL) or large language models (LLMs). Among them, RL agents are efficient for deployment but only perform very few tasks. By contrast, giant LLM agents (often more than 1000B parameters) present strong generalization while demanding enormous computing resources. In this work, we combine their advantages while avoiding the drawbacks by conducting the proposed referee RL on our developed large auto-regressive model (LARM). Specifically, LARM is built upon a lightweight LLM (fewer than 5B parameters) and directly outputs the next action to execute rather than text. We mathematically reveal that classic RL feedbacks vanish in long-horizon embodied exploration and introduce a giant LLM based referee to handle this reward vanishment during training LARM. In this way, LARM learns to complete diverse open-world tasks without human intervention. Especially, LARM successfully harvests enchanted diamond equipment in Minecraft, which demands significantly longer decision-making chains than the highest achievements of prior best methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- DriveGPT4-V2: Harnessing Large Language Model Capabilities for Enhanced Closed-Loop Autonomous DrivingZhenhua Xu, Yan Bai, Yujia Zhang, Zhuoling Li 等CVPR 2025
- Cognitive Predictive Processing: A Human-inspired Framework for Adaptive Exploration in Open-World Reinforcement LearningBoheng Liu, Ziyu Li, Chenghua Duan, Yutian Liu 等NeurIPS 2025
- VistaWise: Building Cost-Effective Agent with Cross-Modal Knowledge Graph for MinecraftHonghao Fu, Junlong Ren, Qi Chai, Deheng Ye 等EMNLP 2025
它引用的顶会 Paper7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra 等ICCV 2019 · 被引用 1,863 次
相关 Paper
- RL-GPT: Integrating Reinforcement Learning and Code-as-policyShaoteng Liu, Haoqi Yuan, Minda Hu, Yanwei Li 等NeurIPS 2024 · 被引用 48 次
- Large Language Models as Generalizable Policies for Embodied TasksAndrew Szot, Max Schwarzer, Harsh Agrawal, Bogdan Mazoure 等ICLR 2024 · 被引用 114 次
- From Multimodal LLMs to Generalist Embodied Agents: Methods and LessonsAndrew Szot, Bogdan Mazoure, Omar Attia, Aleksei Timofeev 等CVPR 2025
- VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning TasksShiduo Zhang, Zhe Xu, Peiju Liu, Xiaopeng Yu 等ICCV 2025 · 被引用 12 次
- ModularAgent: A Task-Aware Modular Framework for Joint Optimization of Multimodal Large Language Models and World ModelsYu-Wei Zhan, Xin Wang, Pengzhe Mao, Tongtong Feng 等CVPR 2026
