Latent Action Learning Requires Supervision in the Presence of Distractors
Alexander Nikulin, Ilya Zisman, Denis Tarasov, Nikita Lyubaykin, Andrei Polubarov, Igor Kiselev, Vladislav Kurenkov
Abstract
Recently, latent action learning, pioneered by Latent Action Policies (LAPO), have shown remarkable pre-training efficiency on observationonly data, offering potential for leveraging vast amounts of video available on the web for embodied AI. However, prior work has focused on distractor-free data, where changes between observations are primarily explained by ground-truth actions. Unfortunately, real-world videos contain action-correlated distractors that may hinder latent action learning. Using Distracting Control Suite (DCS) we empirically investigate the effect of distractors on latent action learning and demonstrate that LAPO struggle in such scenario. We propose LAOM, a simple LAPO modification that improves the quality of latent actions by 8x, as measured by linear probing. Importantly, we show that providing supervision with ground-truth actions, as few as 2.5% of the full dataset, during latent action learning improves downstream performance by 4.2x on average. Our findings suggest that integrating supervision during Latent Action Models (LAM) training is critical in the presence of distractors, challenging the conventional pipeline of first learning LAM and only then decoding from latent to ground-truth actions. Latent Action Learning Requires Supervision in the Presence of Distractors cheetah-run walker-run hopper-hop humanoid-walk Figure 2. Visualization of the environments from the Distracting Control Suite (DCS) used in our work. Top row: without any distractors, identical to the original DeepMind Control Suite. Bottom row: with distractors, which consists of dynamic background videos, agent color change and camera shaking. See Section 3 for additional details.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1892a65e-ff10-4488-a366-daebba09fc4cCited by top-tier papers19
- Motus: A Unified Latent Action World ModelHongzhe Bi, Hengkai Tan, Shenghao Xie, Zeyuan Wang et al.CVPR 2026 · 271 citations
- villa-X: Enhancing Latent Action Modeling in Vision-Language-Action ModelsXiaoyu Chen, Hangxing Wei, Pushi Zhang, Chuheng Zhang et al.ICLR 2026 · 59 citations
- CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot LearningJiange Yang, Yansong Shi, Haoyi Zhu, Mingyu Liu et al.CVPR 2026 · 47 citations
- Learning Latent Action World Models in the WildQuentin Garrido, Tushar Nagarajan, Basile Terver, Nicolas Ballas et al.ICML 2026 · 38 citations
- What Do Latent Action Models Actually Learn?Chuheng Zhang, Tim Pearce, Pushi Zhang, Kaixin Wang et al.NeurIPS 2025 · 35 citations
Builds on26
- Leveraging Procedural Generation to Benchmark Reinforcement LearningKarl Cobbe, Christopher Hesse, Jacob Hilton, John SchulmanICML 2020 · 685 citations
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
- Genie: Generative Interactive EnvironmentsJake Bruce, Michael D. Dennis, Ashley Edwards, Jack Parker-Holder et al.ICML 2024 · 513 citations
- The Pitfalls of Simplicity Bias in Neural NetworksHarshay Shah, Kaustav Tamuly, Aditi Raghunathan, Prateek Jain et al.NeurIPS 2020 · 503 citations
- Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online VideosBowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga et al.NeurIPS 2022 · 458 citations
Related papers
- Object-Centric Latent Action LearningAlbina Klepach, Alexander Nikulin, Ilya Zisman, Denis Tarasov et al.AAAI 2026 · 7 citations
- LAOF: Robust Latent Action Learning with Optical Flow ConstraintsXizhou Bu, Jiexi Lyu, Fulei Sun, Ruichen Yang et al.CVPR 2026 · 10 citations
- Learning to Act without ActionsDominik Schmidt, Minqi JiangICLR 2024 · 98 citations
- AdaWorld: Learning Adaptable World Models with Latent ActionsShenyuan Gao, Siyuan Zhou, Yilun Du, Jun Zhang et al.ICML 2025
- Olaf-World: Orienting Latent Actions for Video World ModelingYuxin Jiang, Yuchao Gu, Ivor Tsang, Mike Zheng ShouICML 2026
