Co-Evolving Latent Action World Models
Yucen Wang, Fengming Zhang, De-Chuan Zhan, Li Zhao, Kaixin Wang, Jiang Bian
Abstract
Adapting pre-trained video generation models into controllable world models via latent actions is a promising step towards creating generalist world models. The dominant paradigm adopts a two-stage approach that trains latent action model (LAM) and the world model separately, resulting in redundant training and limiting their potential for co-adaptation. A conceptually simple and appealing idea is to directly replace the forward dynamic model in LAM with a powerful world model and training them jointly, but it is non-trivial and prone to representational collapse. In this work, we propose CoLA-World, which for the first time successfully realizes this synergistic paradigm, resolving the core challenge in joint learning through a critical warm-up phase that effectively aligns the representations of the from-scratch LAM with the pre-trained world model. This unlocks a co-evolution cycle: the world model acts as a knowledgeable tutor, providing gradients to shape a high-quality LAM, while the LAM offers a more precise and adaptable control interface to the world model. Empirically, CoLA-World matches or outperforms prior two-stage methods in both video simulation quality and downstream visual planning, establishing a robust and efficient new paradigm for the field.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2f5fb1e5-0713-4e55-9b87-d26ac81eb160Cited by top-tier papers4
- DreamDojo: A Real-Time Robot World Model from Large-Scale Human VideosShenyuan Gao, William Liang, Kaiyuan Zheng, Ayaan Malik et al.ICML 2026 · 96 citations
- Learning Latent Action World Models in the WildQuentin Garrido, Tushar Nagarajan, Basile Terver, Nicolas Ballas et al.ICML 2026 · 38 citations
- VideoWorld 2: Learning Transferable Knowledge from Real-world VideosZhongwei Ren, Yunchao Wei, Xiao Yu, Guixun Luo et al.CVPR 2026 · 9 citations
- Multi-view Consistent Latent Action Learning for World Modeling and ControlShenghua Wan, Xiaohai Hu, Xunlan Zhou, lei yuan et al.ICML 2026
Builds on15
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 1,170 citations
- Genie: Generative Interactive EnvironmentsJake Bruce, Michael D. Dennis, Ashley Edwards, Jack Parker-Holder et al.ICML 2024 · 513 citations
- HoloAssist: an Egocentric Human Interaction Dataset for Interactive AI Assistants in the Real WorldXin Wang, Taein Kwon, Mahdi Rad, Bowen Pan et al.ICCV 2023 · 151 citations
- HOI4D: A 4D Egocentric Dataset for Category-Level Human-Object InteractionYunze Liu, Yun Liu, Che Jiang, Kangbo Lyu et al.CVPR 2022 · 126 citations
Related papers
- AdaWorld: Learning Adaptable World Models with Latent ActionsShenyuan Gao, Siyuan Zhou, Yilun Du, Jun Zhang et al.ICML 2025
- Chain of World: World Model Thinking in Latent MotionFuxiang Yang, Donglin Di, Lulu Tang, Xuancheng Zhang et al.CVPR 2026 · 11 citations
- DiLA: Disentangled Latent Action World ModelsTianqiu Zhang, Muyang Lyu, Yufan Zhang, Fang Fang et al.ICML 2026 · 2 citations
- Controlling Large Language Model with Latent ActionChengxing Jia, Ziniu Li, Pengyuan Wang, Yi-Chen Li et al.ICML 2025
- From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot ManipulationYajie Li, Bozhou Zhang, Chun Gu, Zipei Ma et al.ICML 2026 · 2 citations
