Factored Latent Action World Models
Zizhao Wang, Chang Shi, Jiaheng Hu, Kevin Rohling, Roberto Martín-Martín, Amy Zhang, Peter Stone
Abstract
Learning latent actions from action-free video has emerged as a powerful paradigm for scaling up controllable world model learning. Latent actions provide a natural interface for users to iteratively generate and manipulate videos. However, most existing approaches rely on monolithic inverse and forward dynamics models that learn a single latent action to control the entire scene, and therefore struggle in complex environments where multiple entities act simultaneously. This paper introduces Factored Latent Action Model (FLAM), a factored dynamics framework that decomposes the scene into independent factors, each inferring its own latent action and predicting its own next-step factor value. This factorized structure enables more accurate modeling of complex multi-entity dynamics and improves video generation quality in action-free video settings compared to monolithic models. Based on experiments on both simulation and real-world multi-entity datasets, we find that FLAM outperforms prior work in prediction accuracy and representation quality, and facilitates downstream policy learning, demonstrating the benefits of factorized latent action models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on20
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran et al.NeurIPS 2020 · 1,275 citations
- Leveraging Procedural Generation to Benchmark Reinforcement LearningKarl Cobbe, Christopher Hesse, Jacob Hilton, John SchulmanICML 2020 · 685 citations
- Genie: Generative Interactive EnvironmentsJake Bruce, Michael D. Dennis, Ashley Edwards, Jack Parker-Holder et al.ICML 2024 · 513 citations
- Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online VideosBowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga et al.NeurIPS 2022 · 458 citations
- Finite Scalar Quantization: VQ-VAE Made SimpleFabian Mentzer, David Minnen, Eirikur Agustsson, Michael TschannenICLR 2024 · 442 citations
Related papers
- DiLA: Disentangled Latent Action World ModelsTianqiu Zhang, Muyang Lyu, Yufan Zhang, Fang Fang et al.ICML 2026 · 2 citations
- Motus: A Unified Latent Action World ModelHongzhe Bi, Hengkai Tan, Shenghao Xie, Zeyuan Wang et al.CVPR 2026 · 271 citations
- Co-Evolving Latent Action World ModelsYucen Wang, Fengming Zhang, De-Chuan Zhan, Li Zhao et al.ICML 2026 · 12 citations
- From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot ManipulationYajie Li, Bozhou Zhang, Chun Gu, Zipei Ma et al.ICML 2026 · 2 citations
- Disentangled Robot Learning via Separate Forward and Inverse Dynamics PretrainingWenyao Zhang, Bozhou Zhang, Zekun Qi, Wenjun Zeng et al.ICLR 2026 · 18 citations
