RoboOmni: Actions Are Just Another Modality for Vision-Language Models
Dong Wang, Zilong Chen, Jirong Liu, Ziqing Qiao, Xin Xiao, Bingyi Kang, Hongtao Wu, Xiao Ma, Tao Kong, Huaping Liu
摘要
Integrating Vision-Language Models (VLMs) into robotics has facilitated the development of generalizable Vision-Language Action (VLA) policies. However, unified discrete frameworks lag behind decoupled continuous designs due to limitations in action chunking and temporal modeling. To address this, we introduce RoboOmni, a unified multi-modal next-token prediction framework. Challenging the assumption that continuous modeling is essential for high-performance manipulation, RoboOmni demonstrates that actions are just another modality capable of being effectively modeled discretely. At the core of our method is Multi-Token Action Prediction (MTAP), which integrates action chunking directly into the discrete tokenizer. This design resolves temporal modeling bottlenecks and significantly reduces distribution shift between training and inference. By preserving the native VLM training and inference pipeline, RoboOmni naturally benefits from large-scale multimodal co-training and modern decoding optimizations. Extensive evaluations on the CALVIN, SimplerEnv, and real-world platforms confirm that RoboOmni establishes new state-of-the-art performance, significantly outperforming diffusion-based baselines such as . Notably, combining our proposed MTAP with the FAST tokenizer achieves a 94.4% average success rate on CALVIN, while the Bin tokenizer implementation attains a 27 inference speedup compared to OpenVLA.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun 等NeurIPS 2024 · 被引用 1,586 次
- Grounding Multimodal Large Language Models to the WorldZhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao 等ICLR 2024 · 被引用 1,170 次
- Vision-Language Foundation Models as Effective Robot ImitatorsXinghang Li, Minghuan Liu, Hanbo Zhang, Cunjun Yu 等ICLR 2024 · 被引用 375 次
- Unleashing Large-Scale Video Generative Pre-training for Visual Robot ManipulationHongtao Wu, Ya Jing, Chilam Cheang, Guangzeng Chen 等ICLR 2024 · 被引用 309 次
- Better & Faster Large Language Models via Multi-token PredictionFabian Gloeckle, Badr Youbi Idrissi, Baptiste Rozière, David Lopez-Paz 等ICML 2024 · 被引用 286 次
相关 Paper
- FASTer: Toward Powerful and Efficient Autoregressive Vision-Language-Action Models with Learnable Action Tokenizer and Block-wise DecodingYicheng Liu, Shiduo Zhang, Zibin Dong, Baijun Ye 等ICLR 2026
- Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action ModelingHao Li, Shuai Yang, Yilun Chen, Xinyi Chen 等AAAI 2026
- HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action ModelJiaming Liu, Hao Chen, Zhuoyang Liu, Pengju An 等ICLR 2026 · 被引用 216 次
- MoEActok: A MoE-based Action Tokenizer for Vision-Language-Action ModelsChunpu Xu, Zhixuan Liang, Tianshuo Yang, Chi-Min Chan 等CVPR 2026 · 被引用 1 次
- RoboOmni: Proactive Robot Manipulation in Omni-modal ContextSiyin Wang, Jinlan Fu, Feihong Liu, Xinzhe He 等ICLR 2026 · 被引用 9 次
