DiLA: Disentangled Latent Action World Models
Tianqiu Zhang, Muyang Lyu, Yufan Zhang, Fang Fang, Si Wu
摘要
Latent Action Models (LAMs) enable the learning of world models from unlabeled video by inferring abstract actions between consecutive frames. However, LAMs face a fundamental trade-off between action abstraction and generation fidelity. Existing methods typically circumvent this issue by using two-stage training with pre-trained world models or by limiting predictions to optical flow. In this paper, we introduce DiLA , a novel Di sentangled L atent A ction world model that aims to resolve this trade-off via content-structure disentanglement. Our key insight is that disentanglement and latent action learning are co-evolving: the predictive bottleneck inherent in latent action learning serves as a driving force for disentanglement, compelling the model to distill spatial layouts into the structure pathway while offloading visual details to a separate content pathway for generation. This synergy yields a continuous, semantically structured latent action space without compromising generative quality. DiLA achieves superior results in video generation quality, action transfer, visual planning, and manifold interpretability. These findings establish DiLA as a unified framework that simultaneously achieves high-level action abstraction and high-fidelity generation, advancing the frontier of self-supervised world model learning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- Diffusion Transformers with Representation AutoencodersBoyang Zheng, Nanye Ma, Shengbang Tong, Saining XieICLR 2026 · 被引用 288 次
- Motus: A Unified Latent Action World ModelHongzhe Bi, Hengkai Tan, Shenghao Xie, Zeyuan Wang 等CVPR 2026 · 被引用 271 次
- Learning from Reward-Free Offline Data: A Case for Planning with Latent Dynamics ModelsUladzislau Sobal, Wancong Zhang, Kyunghyun Cho, Randall Balestriero 等NeurIPS 2025 · 被引用 109 次
相关 Paper
- Disentangled Robot Learning via Separate Forward and Inverse Dynamics PretrainingWenyao Zhang, Bozhou Zhang, Zekun Qi, Wenjun Zeng 等ICLR 2026 · 被引用 18 次
- Co-Evolving Latent Action World ModelsYucen Wang, Fengming Zhang, De-Chuan Zhan, Li Zhao 等ICML 2026 · 被引用 12 次
- Multi-view Consistent Latent Action Learning for World Modeling and ControlShenghua Wan, Xiaohai Hu, Xunlan Zhou, lei yuan 等ICML 2026
- Olaf-World: Orienting Latent Actions for Video World ModelingYuxin Jiang, Yuchao Gu, Ivor Tsang, Mike Zheng ShouICML 2026
- VDSM: Unsupervised Video Disentanglement With State-Space Modeling and Deep Mixtures of ExpertsMatthew J. Vowels, Necati Cihan Camgöz, Richard BowdenCVPR 2021
