Motus: A Unified Latent Action World Model
Hongzhe Bi, Hengkai Tan, Shenghao Xie, Zeyuan Wang, Shuhe Huang, Haitian Liu, Ruowen Zhao, Yao Feng, Chendong Xiang, Yinze Rong, Hongyan Zhao, Hanyu Liu
摘要
While a general embodied agent must function as a unified system, current methods are built on isolated models for understanding, world modeling, and control. This fragmentation prevents unifying multimodal generative capabilities and hinders learning from large-scale, heterogeneous data. In this paper, we propose Motus, a unified latent action world model that leverages existing general pretrained models and rich, sharable motion information. Motus introduces a Mixture-of-Transformer (MoT) architecture to integrate three experts (i.e., understanding, video generation, and action) and adopts a UniDiffuser-style scheduler to enable flexible switching between different modeling modes (i.e., world models, vision-language-action models, inverse dynamics models, video generation models, and video-action joint prediction models). Motus further leverages the optical flow to learn latent actions and adopts a recipe with three-phase training pipeline and six-layer data pyramid, thereby extracting pixel-level "delta action" and enabling large-scale action pretraining. Experiments show that Motus achieves superior performance against state-of-the-art methods in both simulation (a +15% improvement over X-VLA and a +45% improvement over π 0.5 ) and real-world scenarios(improved by +11 48%), demonstrating unified modeling of all functionalities and priors significantly benefits downstream robotic tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action ModelsBorong Zhang, Jiahao Li, Jiachen Shen, Yuhao Zhang 等ICML 2026 · 被引用 25 次
- VLANeXt: Recipes for Building Strong VLA ModelsXiao-Ming Wu, Bin Fan, Kang Liao, Jian-Jian Jiang 等ICML 2026 · 被引用 10 次
- From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot ManipulationYajie Li, Bozhou Zhang, Chun Gu, Zipei Ma 等ICML 2026 · 被引用 2 次
- From Pixels to Tokens: A Systematic Study of Latent Action Supervision for Vision-Language-Action ModelsYihan Lin, Haoyang Li, Yang Li, Haitao Shen 等ICML 2026 · 被引用 2 次
- DiLA: Disentangled Latent Action World ModelsTianqiu Zhang, Muyang Lyu, Yufan Zhang, Fang Fang 等ICML 2026 · 被引用 2 次
它引用的顶会 Paper28
- Learning Interactive Real-World SimulatorsSherry Yang, Yilun Du, Seyed Kamyar Seyed Ghasemipour, Jonathan Tompson 等ICLR 2024 · 被引用 399 次
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic ManipulationTianxing Chen, Zanxin Chen, Baijun Chen, Zijian Cai 等ICML 2026 · 被引用 394 次
- X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action ModelJinliang Zheng, Jianxiong Li, Zhihao Wang, Dongxiu Liu 等ICLR 2026 · 被引用 335 次
- Zero-Shot Robotic Manipulation with Pre-Trained Image-Editing Diffusion ModelsKevin Black, Mitsuhiko Nakamoto, Pranav Atreya, Homer Rich Walke 等ICLR 2024 · 被引用 284 次
- Show-o2: Improved Native Unified Multimodal ModelsJinheng Xie, Zhenheng Yang, Mike Zheng ShouNeurIPS 2025 · 被引用 261 次
相关 Paper
- Chain of World: World Model Thinking in Latent MotionFuxiang Yang, Donglin Di, Lulu Tang, Xuancheng Zhang 等CVPR 2026 · 被引用 11 次
- VideoVLA: Video Generators Can Be Generalizable Robot ManipulatorsYichao Shen, Fangyun Wei, Zhiying Du, Yaobo Liang 等NeurIPS 2025 · 被引用 73 次
- UniJEPA: Enhancing Robot Policy via Unified Continuous and Discrete Representation LearningJianke Zhang, Yucheng Hu, Yanjiang Guo, Xiaoyu Chen 等ICML 2026
- LLaMo: Scaling Pretrained Language Models for Unified Motion Understanding and Generation with Continuous Autoregressive TokensZekun Li, Sizhe An, Chengcheng Tang, Chuan Guo 等CVPR 2026 · 被引用 12 次
- XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion RepresentationsShichao Fan, Kun Wu, Zhengping Che, Xinhua Wang 等ICML 2026 · 被引用 16 次
