HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model
Jiaming Liu, Hao Chen, Zhuoyang Liu, Pengju An, Renrui Zhang, Chenyang Gu, Xiaoqi Li, Ziyu Guo, Sixiang Chen, Mengzhen Liu, Chengkai Hou, Mengdi Zhao
摘要
A fundamental objective of manipulation policy design is to endow robots to comprehend human instructions, reason about scene cues, and execute generalized actions in dynamic environments. Recent autoregressive vision-language-action (VLA) methods inherit common-sense reasoning capabilities from vision-language models (VLMs) for next action-token prediction. However, these methods quantize actions into discrete bins, which disrupts the continuity required for precise control. In contrast, existing diffusion-based VLA methods incorporate an additional diffusion head to predict continuous actions solely conditioned on feature representations extracted by the VLM, without fully leveraging the VLM's pretrained reasoning capabilities through token-level generation. To address these limitations, we introduce HybridVLA, a unified framework that absorbs the continuous nature of diffusion-based actions and the contextual reasoning of autoregression within a single large language model. To mitigate interference between the two generation paradigms, we propose a collaborative training recipe that seamlessly incorporates diffusion denoising into the next-token prediction process. With this recipe, we find these two action prediction methods not only reinforce each other but also exhibit varying strength across different tasks. Therefore, we design a collaborative action ensemble mechanism that adaptively fuses both predictions, leading to more robust control. HybridVLA outperforms previous state-of-the-art VLA methods by 14% and 19% in mean success rate on simulation and real-world tasks, respectively, while demonstrating stable manipulation in unseen configurations. On the one hand, autoregressive VLA methods [9, 11, 10, 15] emulate the reasoning paradigm of VLMs for next token prediction, effectively leveraging their large-scale pretrained knowledge and Preprint. Under review.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper39
- DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World KnowledgeWenyao Zhang, Hongsi Liu, Zekun Qi, Yunnan Wang 等NeurIPS 2025 · 被引用 244 次
- T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoTDongzhi Jiang, Ziyu Guo, Renrui Zhang, Zhuofan Zong 等NeurIPS 2025 · 被引用 181 次
- Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize BetterDanny Driess, Jost Tobias Springenberg, Brian Ichter, Lili Yu 等NeurIPS 2025 · 被引用 162 次
- ForceVLA: Enhancing VLA Models with a Force-aware MoE for Contact-rich ManipulationJiawen Yu, Hairuo Liu, Qiaojun Yu, Jieji Ren 等NeurIPS 2025 · 被引用 150 次
- CogVLA: Cognition-Aligned Vision-Language-Action Models via Instruction-Driven Routing & SparsificationWei Li, Renshan Zhang, Rui Shao, Jie He 等NeurIPS 2025 · 被引用 87 次
它引用的顶会 Paper30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
相关 Paper
- CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action ModelsQingqing Zhao, Yao Lu, Moo Jin Kim, Zipeng Fu 等CVPR 2025
- EnsembleVLA: Ensemble Learning for Vision-Language Action ModelsMingchen Song, Xiang Deng, Jie Wei, Dongmei Jiang 等ICML 2026
- Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action PoliciesZhixuan Liang, Yizhuo Li, Tianshuo Yang, CHENGYUE WU 等ICML 2026 · 被引用 86 次
- VideoVLA: Video Generators Can Be Generalizable Robot ManipulatorsYichao Shen, Fangyun Wei, Zhiying Du, Yaobo Liang 等NeurIPS 2025 · 被引用 73 次
- Vision-Language-Action Instruction Tuning: From Understanding to ManipulationShuai Yang, Hao Li, Bin Wang, Yilun Chen 等ICLR 2026 · 被引用 50 次
