What Do Latent Action Models Actually Learn?
Chuheng Zhang, Tim Pearce, Pushi Zhang, Kaixin Wang, Xiaoyu Chen, Wei Shen, Li Zhao, Jiang Bian
Abstract
Latent action models (LAMs) aim to learn action-relevant changes from unlabeled videos by compressing changes between frames as latents. However, differences between video frames can be caused by controllable changes as well as exogenous noise, leading to an important concern -do latents capture the changes caused by actions or irrelevant noise? This paper studies this issue analytically, presenting a linear model that encapsulates the essence of LAM learning, while being tractable. This provides several insights, including connections between LAM and principal component analysis (PCA), desiderata of the data-generating policy, and justification of strategies to encourage learning controllable changes using data augmentation, data cleaning, and auxiliary action-prediction. These findings are validated through numerical simulations, as well as experiments in more realistic settings. This investigation is the first to rigorously investigate how the structure of observations, actions, and noise influence LAM learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eb5452bf-5603-48d8-a93a-19271af2f98aCited by top-tier papers8
- Motus: A Unified Latent Action World ModelHongzhe Bi, Hengkai Tan, Shenghao Xie, Zeyuan Wang et al.CVPR 2026 · 271 citations
- DreamDojo: A Real-Time Robot World Model from Large-Scale Human VideosShenyuan Gao, William Liang, Kaiyuan Zheng, Ayaan Malik et al.ICML 2026 · 96 citations
- villa-X: Enhancing Latent Action Modeling in Vision-Language-Action ModelsXiaoyu Chen, Hangxing Wei, Pushi Zhang, Chuheng Zhang et al.ICLR 2026 · 59 citations
- World Guidance: World Modeling in Condition Space for Action GenerationYue Su, Sijin Chen, Haixin Shi, Mingyu Liu et al.ICML 2026 · 26 citations
- LAOF: Robust Latent Action Learning with Optical Flow ConstraintsXizhou Bu, Jiexi Lyu, Fulei Sun, Ruichen Yang et al.CVPR 2026 · 10 citations
Builds on15
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
- Genie: Generative Interactive EnvironmentsJake Bruce, Michael D. Dennis, Ashley Edwards, Jack Parker-Holder et al.ICML 2024 · 513 citations
- InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and GenerationYi Wang, Yinan He, Yizhuo Li, Kunchang Li et al.ICLR 2024 · 467 citations
- Learning to Act without ActionsDominik Schmidt, Minqi JiangICLR 2024 · 98 citations
Related papers
- Factored Latent Action World ModelsZizhao Wang, Chang Shi, Jiaheng Hu, Kevin Rohling et al.ICML 2026 · 4 citations
- DiLA: Disentangled Latent Action World ModelsTianqiu Zhang, Muyang Lyu, Yufan Zhang, Fang Fang et al.ICML 2026 · 2 citations
- LARA: Latent Action Representation Alignment for Vision-Language-Action ModelsMengya Liu, Baoxiong Jia, Jiangyong Huang, Jingze Zhang et al.ICML 2026 · 3 citations
- Learning Latent Action World Models in the WildQuentin Garrido, Tushar Nagarajan, Basile Terver, Nicolas Ballas et al.ICML 2026 · 38 citations
- Latent Action Learning Requires Supervision in the Presence of DistractorsAlexander Nikulin, Ilya Zisman, Denis Tarasov, Nikita Lyubaykin et al.ICML 2025
