From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons
Andrew Szot, Bogdan Mazoure, Omar Attia, Aleksei Timofeev, Harsh Agrawal, R. Devon Hjelm, Zhe Gan, Zsolt Kira, Alexander Toshev
摘要
We examine the capability of Multimodal Large Language Models (MLLMs) to tackle diverse domains that extend beyond the traditional language and vision tasks these models are typically trained on. Specifically, our focus lies in areas such as Embodied AI, Games, UI Control, and Planning. To this end, we introduce a process of adapting an MLLM to a Generalist Embodied Agent (GEA). GEA is a single unified model capable of grounding itself across these varied domains through a multi-embodiment action tokenizer. GEA is trained with supervised learning on a large dataset of embodied experiences and with online RL in interactive simulators. We explore the data and algorithmic choices necessary to develop such a model. Our findings reveal the importance of training with cross-domain data and online RL for building generalist agents. The final GEA model achieves strong generalization performance to unseen tasks across diverse benchmarks compared to other generalist models and benchmark-specific approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent PlanningChi-Pin Huang, Yueh-Hua Wu, Min-Hung Chen, Yu-Chiang Frank Wang 等NeurIPS 2025 · 被引用 179 次
- Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize BetterDanny Driess, Jost Tobias Springenberg, Brian Ichter, Lili Yu 等NeurIPS 2025 · 被引用 162 次
- BEAST: Efficient Tokenization of B-Splines Encoded Action Sequences for Imitation LearningHongyi Zhou, Weiran Liao, Xi Huang, Yucheng Tang 等NeurIPS 2025 · 被引用 32 次
- From Indoor to Open World: Revealing the Spatial Reasoning Gap in MLLMsMingrui Wu, Zhaozhi Wang, Fangjinhua Wang, Jiaolong Yang 等CVPR 2026 · 被引用 11 次
- MAPS: Preserving Vision-Language Representations via Module-Wise Proximity Scheduling for Better Vision-Language-Action GeneralizationChengyue Huang, Mellon M. Zhang, Robert Azarcon, Glen Chou 等CVPR 2026 · 被引用 8 次
它引用的顶会 Paper38
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch 等ICML 2023 · 被引用 2,601 次
相关 Paper
- Grounding Multimodal Large Language Models in ActionsAndrew Szot, Bogdan Mazoure, Harsh Agrawal, R. Devon Hjelm 等NeurIPS 2024 · 被引用 43 次
- ModularAgent: A Task-Aware Modular Framework for Joint Optimization of Multimodal Large Language Models and World ModelsYu-Wei Zhan, Xin Wang, Pengzhe Mao, Tongtong Feng 等CVPR 2026
- An Embodied Generalist Agent in 3D WorldJiangyong Huang, Silong Yong, Xiaojian Ma, Xiongkun Linghu 等ICML 2024 · 被引用 361 次
- Multitask Multimodal Prompted Training for Interactive Embodied Task CompletionGeorgios Pantazopoulos, Malvina Nikandrou, Amit Parekh, Bhathiya Hemanthage 等EMNLP 2023 · 被引用 1 次
- Large Language Models as Generalizable Policies for Embodied TasksAndrew Szot, Max Schwarzer, Harsh Agrawal, Bogdan Mazoure 等ICLR 2024 · 被引用 114 次
