AdaWorld: Learning Adaptable World Models with Latent Actions
Shenyuan Gao, Siyuan Zhou, Yilun Du, Jun Zhang, Chuang Gan
Abstract
World models aim to learn action-controlled future prediction and have proven essential for the development of intelligent agents. However, most existing world models rely heavily on substantial action-labeled data and costly training, making it challenging to adapt to novel environments with heterogeneous actions through limited interactions. This limitation can hinder their applicability across broader domains. To overcome this limitation, we propose AdaWorld, an innovative world model learning approach that enables efficient adaptation. The key idea is to incorporate action information during the pretraining of world models. This is achieved by extracting latent actions from videos in a self-supervised manner, capturing the most critical transitions between frames. We then develop an autoregressive world model that conditions on these latent actions. This learning paradigm enables highly adaptable world models, facilitating efficient transfer and learning of new actions even with limited interactions and finetuning. Our comprehensive experiments across multiple environments demonstrate that AdaWorld achieves superior performance in both simulation quality and visual planning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext badc3216-74b8-49f4-9d6a-4aa1db1f2652Cited by top-tier papers6
- From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot ManipulationYajie Li, Bozhou Zhang, Chun Gu, Zipei Ma et al.ICML 2026 · 2 citations
- Learning to Theorize the World from ObservationDoojin Baek, Gyubin Lee, Junyeob Baek, Hosung Lee et al.ICML 2026 · 1 citation
- Motion Dynamics Learning for Few-Shot Embodied AdaptationSibo He, Weiying Xie, Daixun Li, Junhao Zhong et al.ICML 2026
- Multi-view Consistent Latent Action Learning for World Modeling and ControlShenghua Wan, Xiaohai Hu, Xunlan Zhou, lei yuan et al.ICML 2026
- Cross-Embodiment Robot Foundation World Models with Latent ActionsHuang Huang, Sriram Yenamandra, Arjun Majumdar, Elie Aljalbout et al.ICML 2026
Builds on44
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 3,959 citations
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra et al.ICCV 2019 · 1,863 citations
- Model Based Reinforcement Learning for AtariLukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski et al.ICLR 2020 · 969 citations
Related papers
- Learning Latent Action World Models in the WildQuentin Garrido, Tushar Nagarajan, Basile Terver, Nicolas Ballas et al.ICML 2026 · 38 citations
- Vid2World: Crafting Video Diffusion Models to Interactive World ModelsSiqiao Huang, Jialong Wu, Qixing Zhou, Shangchen Miao et al.ICLR 2026 · 68 citations
- Co-Evolving Latent Action World ModelsYucen Wang, Fengming Zhang, De-Chuan Zhan, Li Zhao et al.ICML 2026 · 12 citations
- Pre-Trained Video Generative Models as World SimulatorsHaoran He, Yang Zhang, Liang Lin, Zhongwen Xu et al.AAAI 2026 · 32 citations
- Olaf-World: Orienting Latent Actions for Video World ModelingYuxin Jiang, Yuchao Gu, Ivor Tsang, Mike Zheng ShouICML 2026
