VIMA: Robot Manipulation with Multimodal Prompts
Yunfan Jiang, Agrim Gupta, Zichen Zhang, Guanzhi Wang, Yongqiang Dou, Yanjun Chen, Li Fei-Fei, Anima Anandkumar, Yuke Zhu, Linxi Fan
摘要
Prompt-based learning has emerged as a successful paradigm in natural language processing, where a single general-purpose language model can be instructed to perform any task specified by input prompts. Yet task specification in robotics comes in various forms, such as imitating one-shot demonstrations, following language instructions, and reaching visual goals. They are often considered different tasks and tackled by specialized models. We show that a wide spectrum of robot manipulation tasks can be expressed with multimodal prompts, interleaving textual and visual tokens. Accordingly, we develop a new simulation benchmark that consists of thousands of procedurally-generated tabletop tasks with multimodal prompts, 600K+ expert trajectories for imitation learning, and a four-level evaluation protocol for systematic generalization. We design a transformer-based robot agent, VIMA, that processes these prompts and outputs motor actions autoregressively. VIMA features a recipe that achieves strong model scalability and data efficiency. It outperforms alternative designs in the hardest zero-shot generalization setting by up to task success rate given the same training data. With less training data, VIMA still performs better than the best competing variant. Code and video demos are available at https://vimalabs.github.io/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper26
- ForceVLA: Enhancing VLA Models with a Force-aware MoE for Contact-rich ManipulationJiawen Yu, Hairuo Liu, Qiaojun Yu, Jieji Ren 等NeurIPS 2025 · 被引用 150 次
- True Knowledge Comes from Practice: Aligning Large Language Models with Embodied Environments via Reinforcement LearningWeihao Tan, Wentao Zhang, Shanqi Liu, Longtao Zheng 等ICLR 2024 · 被引用 33 次
- ForceVLA2: Unleashing Hybrid Force-Position Control with Force Awareness for Contact-Rich ManipulationYang Li, Zhaxizhuoma, Hongru Jiang, Junjie Xia 等CVPR 2026 · 被引用 31 次
- AdvEDM: Fine-grained Adversarial Attack against VLM-based Embodied AgentsYichen Wang, Hangtao Zhang, Hewen Pan, Ziqi Zhou 等NeurIPS 2025 · 被引用 27 次
- Embodied Multi-Modal Agent trained by an LLM from a Parallel TextWorldYijun Yang, Tianyi Zhou, Kanxue Li, Dapeng Tao 等CVPR 2024 · 被引用 23 次
它引用的顶会 Paper42
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch 等ICML 2023 · 被引用 2,601 次
相关 Paper
- Mastering Robot Manipulation with Multimodal Prompts through Pretraining and Multi-task Fine-tuningJiachen Li, Qiaozi Gao, Michael Johnston, Xiaofeng Gao 等ICML 2024 · 被引用 18 次
- PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMsSoroush Nasiriany, Fei Xia, Wenhao Yu, Ted Xiao 等ICML 2024 · 被引用 212 次
- GenSim: Generating Robotic Simulation Tasks via Large Language ModelsLirui Wang, Yiyang Ling, Zhecheng Yuan, Mohit Shridhar 等ICLR 2024 · 被引用 143 次
- Vision-Language Models are Zero-Shot Reward Models for Reinforcement LearningJuan Rocamonde, Victoriano Montesinos, Elvis Nava, Ethan Perez 等ICLR 2024 · 被引用 154 次
- PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action ModelWenqi Liang, Gan Sun, Yao He, Jiahua Dong 等ICLR 2026 · 被引用 20 次
