Multi-view Consistent Latent Action Learning for World Modeling and Control
Shenghua Wan, Xiaohai Hu, Xunlan Zhou, lei yuan, Le Gan, De-Chuan Zhan
摘要
The scalability of world models is currently bottlenecked by the scarcity of action annotations. While self-supervised latent action learning offers a potential solution, existing single-view paradigms—relying on information bottlenecks or Vector Quantization (VQ)—often conflate superficial 2D pixel displacements with the underlying physical-spatial dynamics of an action. Consequently, these methods remain highly susceptible to view-dependent noise, such as camera shake. We introduce MuCoLA ( Mu lti-view Co nsistent L atent A ction learning), a framework that learns robust, view-invariant action representations by enforcing semantic consistency across synchronized video streams. MuCoLA utilizes a Student-Teacher network with DINO-style self-distillation to align action distributions across viewpoints, effectively filtering high-frequency visual noise while preserving motion semantics. Theoretical analysis reveals that our multi-view objective functions as a spectral filter, isolating agent dynamics from environmental nuisances. Empirically, MuCoLA significantly outperforms baselines in action regression, video reconstruction, and downstream visual control tasks. Furthermore, we demonstrate that MuCoLA exhibits favorable scaling properties with respect to model capacity and data volume, paving the way for large-scale action-free world modeling.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Genie: Generative Interactive EnvironmentsJake Bruce, Michael D. Dennis, Ashley Edwards, Jack Parker-Holder 等ICML 2024 · 被引用 513 次
- Multi-View Masked World Models for Visual Robotic ManipulationYounggyo Seo, Junsu Kim, Stephen James, Kimin Lee 等ICML 2023 · 被引用 99 次
- Learning to Act without ActionsDominik Schmidt, Minqi JiangICLR 2024 · 被引用 98 次
相关 Paper
- DiLA: Disentangled Latent Action World ModelsTianqiu Zhang, Muyang Lyu, Yufan Zhang, Fang Fang 等ICML 2026 · 被引用 2 次
- Olaf-World: Orienting Latent Actions for Video World ModelingYuxin Jiang, Yuchao Gu, Ivor Tsang, Mike Zheng ShouICML 2026
- Learning Latent Action World Models in the WildQuentin Garrido, Tushar Nagarajan, Basile Terver, Nicolas Ballas 等ICML 2026 · 被引用 38 次
- Motus: A Unified Latent Action World ModelHongzhe Bi, Hengkai Tan, Shenghao Xie, Zeyuan Wang 等CVPR 2026 · 被引用 271 次
- Learning to Act Robustly with View-Invariant Latent ActionsYoungjoon Jeong, Junha Chun, Taesup KimCVPR 2026 · 被引用 5 次
