Actions as Language: Fine-Tuning VLMs into VLAs Without Catastrophic Forgetting
Asher J. Hancock, Xindi Wu, Lihan Zha, Olga Russakovsky, Anirudha Majumdar
摘要
Fine-tuning vision-language models (VLMs) on robot teleoperation data to create vision-languageaction (VLA) models is a promising paradigm for training generalist policies, but it suffers from a fundamental tradeoff: learning to produce actions often diminishes the VLM's foundational reasoning and multimodal understanding, hindering generalization to novel scenarios, instruction following, and semantic understanding. We argue that this catastrophic forgetting is due to a distribution mismatch between the VLM's internet-scale pretraining corpus and the robotics fine-tuning data. Inspired by this observation, we introduce VLM2VLA: a VLA training paradigm that first resolves this mismatch at the data level by representing low-level actions with natural language. This alignment makes it possible to train VLAs solely with Low-Rank Adaptation (LoRA), thereby minimally modifying the VLM backbone and averting catastrophic forgetting. As a result, the VLM can be fine-tuned on robot teleoperation data without fundamentally altering the underlying architecture and without expensive co-training on internet-scale VLM datasets. Through extensive Visual Question Answering (VQA) studies and over 800 real-world robotics experiments, we demonstrate that VLM2VLA preserves the VLM's core capabilities, enabling zero-shot generalization to novel tasks that require open-world semantic reasoning and multilingual instruction following. Website with additional information, videos, and code: https://vlm2vla.github.io/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- MergeVLA: Cross-Skill Model Merging Toward a Generalist Vision-Language-Action AgentYuxia Fu, Zhizhen Zhang, Yuqi Zhang, Zijian Wang 等CVPR 2026 · 被引用 21 次
- LangForce: Bayesian Decomposition of Vision Language Action Models via Latent Action QueriesShijie Lian, Bin Yu, Xiaopeng LIN, Laurence Yang 等ICML 2026 · 被引用 17 次
- MAPS: Preserving Vision-Language Representations via Module-Wise Proximity Scheduling for Better Vision-Language-Action GeneralizationChengyue Huang, Mellon M. Zhang, Robert Azarcon, Glen Chou 等CVPR 2026 · 被引用 8 次
- FASTer: Toward Powerful and Efficient Autoregressive Vision-Language-Action Models with Learnable Action Tokenizer and Block-wise DecodingYicheng Liu, Shiduo Zhang, Zibin Dong, Baijun Ye 等ICLR 2026
- Model-Based Imaginative Planning for Embodied AgentsJunru Song, Hengzhe Jin, Yucong Huang, Tingsong Jiang 等ACL 2026
它引用的顶会 Paper14
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language ModelsSiddharth Karamcheti, Suraj Nair, Ashwin Balakrishna, Percy Liang 等ICML 2024 · 被引用 306 次
- 3D-VLA: A 3D Vision-Language-Action Generative World ModelHaoyu Zhen, Xiaowen Qiu, Peihao Chen, Jincheng Yang 等ICML 2024 · 被引用 303 次
- Generalized Planning in PDDL Domains with Pretrained Large Language ModelsTom Silver, Soham Dan, Kavitha Srinivas, Joshua B. Tenenbaum 等AAAI 2024 · 被引用 194 次
相关 Paper
- Vision-Language-Action Instruction Tuning: From Understanding to ManipulationShuai Yang, Hao Li, Bin Wang, Yilun Chen 等ICLR 2026 · 被引用 50 次
- HMVLM: Human Motion-Vision-Language Model via MoE LoRALei Hu, Yongjing Ye, Shihong XiaNeurIPS 2025 · 被引用 1 次
- SMoLoRa: Exploring and Defying Dual Catastrophic Forgetting in Continual Visual Instruction TuningZiqi Wang, Chang Che, Qi Wang, Yangyang Li 等ICCV 2025 · 被引用 4 次
- OTTER: A Vision-Language-Action Model with Text-Aware Visual Feature ExtractionHuang Huang, Fangchen Liu, Letian Fu, Tingfan Wu 等ICML 2025
- FOCA: Future-Oriented Conditioning for Data-Efficient Vision-Language-Action AdaptationDuc Nguyen, Nghiem Diep, Binh Nguyen Gia, Trong-Bao Ho 等ICML 2026 · 被引用 3 次
