VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers
Yating Wang, Haoyi Zhu, Mingyu Liu, Jiange Yang, Haoshu Fang, Tong He
摘要
In this paper, we introduce an innovative vector quantization based action tokenizer built upon the largest-scale action trajectory dataset to date, leveraging over 100 times more data than previous approaches. This extensive dataset enables our tokenizer to capture rich spatiotemporal dynamics, resulting in a model that not only accelerates inference but also generates smoother and more coherent action outputs. Once trained, the tokenizer can be seamlessly adapted to a wide range of downstream tasks in a zero-shot manner, from short-horizon reactive behaviors to long-horizon planning. A key finding of our work is that the domain gap between synthetic and real action trajectories is marginal, allowing us to effectively utilize a vast amount of synthetic data during training without compromising real-world performance. To validate our approach, we conducted extensive experiments in both simulated environments and on real robotic platforms. The results demonstrate that as the volume of synthetic trajectory data increases, the performance of our tokenizer on downstream tasks improves significantly-most notably, achieving up to a 30% higher success rate on two real-world tasks in long-horizon scenarios. These findings highlight the potential of our action tokenizer as a robust and scalable solution for real-time embodied intelligence systems, paving the way for more efficient and reliable robotic control in diverse application domains.Project website: https://xiaoxiao0406.github.io/vqvla.github.io
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot LearningJiange Yang, Yansong Shi, Haoyi Zhu, Mingyu Liu 等CVPR 2026 · 被引用 47 次
- AtomicVLA: Unlocking the Potential of Atomic Skill Learning in RobotsLikui Zhang, Tao Tang, Zhihao Zhan, Xiuwei Chen 等CVPR 2026 · 被引用 18 次
- StaMo: Unsupervised Learning of Generalizable Robot Motion from Compact State RepresentationMingyu Liu, Jiuhe Shu, Hui Chen, Zeju Li 等CVPR 2026 · 被引用 14 次
- ODYSSEY: Open-World Quadrupeds Exploration and Manipulation for Long-Horizon TasksKaijun Wang, Liqin Lu, Mingyu Liu, Jianuo Jiang 等AAAI 2026 · 被引用 6 次
- Characterizing Vision-Language-Action Models across XPUs: Constraints and Acceleration for On-Robot DeploymentKaijun Zhou, Qiwei Chen, Da Peng, Zhiyang Li 等ICML 2026 · 被引用 4 次
它引用的顶会 Paper17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch 等ICML 2023 · 被引用 2,601 次
- 3D-VLA: A 3D Vision-Language-Action Generative World ModelHaoyu Zhen, Xiaowen Qiu, Peihao Chen, Jincheng Yang 等ICML 2024 · 被引用 303 次
- Behavior Generation with Latent ActionsSeungjae Lee, Yibin Wang, Haritheja Etukuru, H. Jin Kim 等ICML 2024 · 被引用 154 次
相关 Paper
- FASTer: Toward Powerful and Efficient Autoregressive Vision-Language-Action Models with Learnable Action Tokenizer and Block-wise DecodingYicheng Liu, Shiduo Zhang, Zibin Dong, Baijun Ye 等ICLR 2026
- MoEActok: A MoE-based Action Tokenizer for Vision-Language-Action ModelsChunpu Xu, Zhixuan Liang, Tianshuo Yang, Chi-Min Chan 等CVPR 2026 · 被引用 1 次
- BEAST: Efficient Tokenization of B-Splines Encoded Action Sequences for Imitation LearningHongyi Zhou, Weiran Liao, Xi Huang, Yucheng Tang 等NeurIPS 2025 · 被引用 32 次
- InternData-A1: Pioneering High-Fidelity Synthetic Data for Pre-training Generalist PolicyYang Tian, Yuyin Yang, Yiman Xie, Zetao Cai 等CVPR 2026 · 被引用 64 次
- GPC: Large-Scale Generative Pretraining for Transferable Motor ControlYi Shi, Yifeng Jiang, Chen Tessler, Xue Bin PengSIGGRAPH 2026
