Cross-Embodiment Robot Foundation World Models with Latent Actions
Huang Huang, Sriram Yenamandra, Arjun Majumdar, Elie Aljalbout, Tushar Nagarajan, Jimmy Yang, Akshara Rai, Michael Rabbat, Li Fei-Fei, Jiajun Wu, Tingfan Wu, Franziska Meier
Abstract
The diversity of robot embodiments and action spaces makes it challenging to build robot world models that generalize across different embodiments. We introduce the Latent Action-Conditioned Robot World Model (LAC-WM), which operates within a learned unified latent action space shared across diverse embodiments. This unified action space improves the world model’s performance when adapted to previously unseen robot embodiments. We compare LAC-WM with an Explicit Action-Conditioned World Model (EAC-WM), which conditions on explicit motion labels. Our results shows that explicit action conditioning leads to disjoint action representations across embodiments, limiting downstream performance when adapting to new robots. We evaluate both models on dexterous manipulation tasks and a modified LIBERO benchmark. LAC-WM improves downstream performance over EAC-WM by up to 46.7% on dexterous manipulation and 11.7% on LIBERO. Crucially, the unified latent action space allows LAC-WM’s downstream performance to scale positively with the number of embodiments used during pretraining. In contrast, the disjoint action space in EAC-WM leads to decreased performance as the number of pretraining embodiments increases. These results highlights the importance of a unified action space for efficient cross-embodiment learning, addressing a key challenge in robotics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on15
- Genie: Generative Interactive EnvironmentsJake Bruce, Michael D. Dennis, Ashley Edwards, Jack Parker-Holder et al.ICML 2024 · 513 citations
- Learning Interactive Real-World SimulatorsSherry Yang, Yilun Du, Seyed Kamyar Seyed Ghasemipour, Jonathan Tompson et al.ICLR 2024 · 399 citations
- Learning to Act without ActionsDominik Schmidt, Minqi JiangICLR 2024 · 98 citations
- DynaMo: In-Domain Dynamics Pretraining for Visuo-Motor ControlZichen Jeff Cui, Hengkai Pan, Aadhithya Iyer, Siddhant Haldar et al.NeurIPS 2024 · 61 citations
- What Do Latent Action Models Actually Learn?Chuheng Zhang, Tim Pearce, Pushi Zhang, Kaixin Wang et al.NeurIPS 2025 · 35 citations
Related papers
- Cross-Hand Latent Representation for Vision-Language-Action ModelsGuangqi Jiang, Yutong Liang, Jianglong Ye, Jia-Yang Huang et al.CVPR 2026 · 14 citations
- Learning a Unified Latent Action Space from Videos with Action-centric Cycle ConsistencyGuangyan Chen, Qi Shao, Te Cui, Zichen Zhou et al.CVPR 2026
- Learning Latent Action World Models in the WildQuentin Garrido, Tushar Nagarajan, Basile Terver, Nicolas Ballas et al.ICML 2026 · 38 citations
- XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion RepresentationsShichao Fan, Kun Wu, Zhengping Che, Xinhua Wang et al.ICML 2026 · 16 citations
- Grounding Multimodal Large Language Models in ActionsAndrew Szot, Bogdan Mazoure, Harsh Agrawal, R. Devon Hjelm et al.NeurIPS 2024 · 43 citations
