From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons
Andrew Szot, Bogdan Mazoure, Omar Attia, Aleksei Timofeev, Harsh Agrawal, R. Devon Hjelm, Zhe Gan, Zsolt Kira, Alexander Toshev
Abstract
We examine the capability of Multimodal Large Language Models (MLLMs) to tackle diverse domains that extend beyond the traditional language and vision tasks these models are typically trained on. Specifically, our focus lies in areas such as Embodied AI, Games, UI Control, and Planning. To this end, we introduce a process of adapting an MLLM to a Generalist Embodied Agent (GEA). GEA is a single unified model capable of grounding itself across these varied domains through a multi-embodiment action tokenizer. GEA is trained with supervised learning on a large dataset of embodied experiences and with online RL in interactive simulators. We explore the data and algorithmic choices necessary to develop such a model. Our findings reveal the importance of training with cross-domain data and online RL for building generalist agents. The final GEA model achieves strong generalization performance to unseen tasks across diverse benchmarks compared to other generalist models and benchmark-specific approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 021a531e-ea9d-4a93-a35a-4e94ea8ae607Cited by top-tier papers13
- ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent PlanningChi-Pin Huang, Yueh-Hua Wu, Min-Hung Chen, Yu-Chiang Frank Wang et al.NeurIPS 2025 · 179 citations
- Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize BetterDanny Driess, Jost Tobias Springenberg, Brian Ichter, Lili Yu et al.NeurIPS 2025 · 162 citations
- BEAST: Efficient Tokenization of B-Splines Encoded Action Sequences for Imitation LearningHongyi Zhou, Weiran Liao, Xi Huang, Yucheng Tang et al.NeurIPS 2025 · 32 citations
- From Indoor to Open World: Revealing the Spatial Reasoning Gap in MLLMsMingrui Wu, Zhaozhi Wang, Fangjinhua Wang, Jiaolong Yang et al.CVPR 2026 · 11 citations
- MAPS: Preserving Vision-Language Representations via Module-Wise Proximity Scheduling for Better Vision-Language-Action GeneralizationChengyue Huang, Mellon M. Zhang, Robert Azarcon, Glen Chou et al.CVPR 2026 · 8 citations
Builds on38
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch et al.ICML 2023 · 2,601 citations
Related papers
- Grounding Multimodal Large Language Models in ActionsAndrew Szot, Bogdan Mazoure, Harsh Agrawal, R. Devon Hjelm et al.NeurIPS 2024 · 43 citations
- ModularAgent: A Task-Aware Modular Framework for Joint Optimization of Multimodal Large Language Models and World ModelsYu-Wei Zhan, Xin Wang, Pengzhe Mao, Tongtong Feng et al.CVPR 2026
- An Embodied Generalist Agent in 3D WorldJiangyong Huang, Silong Yong, Xiaojian Ma, Xiongkun Linghu et al.ICML 2024 · 361 citations
- Multitask Multimodal Prompted Training for Interactive Embodied Task CompletionGeorgios Pantazopoulos, Malvina Nikandrou, Amit Parekh, Bhathiya Hemanthage et al.EMNLP 2023 · 1 citation
- Large Language Models as Generalizable Policies for Embodied TasksAndrew Szot, Max Schwarzer, Harsh Agrawal, Bogdan Mazoure et al.ICLR 2024 · 114 citations
