Learning World Models with Identifiable Factorization
Yuren Liu, Biwei Huang, Zhengmao Zhu, Hong-Long Tian, Mingming Gong, Yang Yu, Kun Zhang
摘要
Extracting a stable and compact representation of the environment is crucial for efficient reinforcement learning in high-dimensional, noisy, and non-stationary environments. Different categories of information coexist in such environments -- how to effectively extract and disentangle these information remains a challenging problem. In this paper, we propose IFactor, a general framework to model four distinct categories of latent state variables that capture various aspects of information within the RL system, based on their interactions with actions and rewards. Our analysis establishes block-wise identifiability of these latent variables, which not only provides a stable and compact representation but also discloses that all reward-relevant factors are significant for policy learning. We further present a practical approach to learning the world model with identifiable blocks, ensuring the removal of redundants but retaining minimal and sufficient information for policy optimization. Experiments in synthetic worlds demonstrate that our method accurately identifies the ground-truth latent variables, substantiating our theoretical findings. Moreover, experiments in variants of the DeepMind Control Suite and RoboDesk showcase the superior performance of our approach over baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Focus On What Matters: Separated Models For Visual-Based RL GeneralizationDi Zhang, Bowen Lv, Hai Zhang, Feifan Yang 等NeurIPS 2024 · 被引用 14 次
- NeuralOS: Towards Simulating Operating Systems via Neural Generative ModelsLuke Rivard, Sun Sun, Hongyu Guo, Wenhu Chen 等ICLR 2026 · 被引用 13 次
- Leveraging Separated World Model for Exploration in Visually Distracted EnvironmentsKaichen Huang, Shenghua Wan, Minghao Shao, Hai-Hang Sun 等NeurIPS 2024 · 被引用 5 次
- BECAUSE: Bilinear Causal Representation for Generalizable Offline Model-based Reinforcement LearningHaohong Lin, Wenhao Ding, Jian Chen, Laixi Shi 等NeurIPS 2024 · 被引用 5 次
- DDP-WM: Disentangled Dynamics Prediction for Efficient World ModelsShicheng Yin, Kaixuan Yin, Weixing Chen, Yang Liu 等ICML 2026 · 被引用 3 次
它引用的顶会 Paper20
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 被引用 1,261 次
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 被引用 1,170 次
- Self-Supervised Learning with Data Augmentations Provably Isolates Content from StyleJulius von Kügelgen, Yash Sharma, Luigi Gresele, Wieland Brendel 等NeurIPS 2021 · 被引用 421 次
- Scalable Methods for Computing State Similarity in Deterministic Markov Decision ProcessesPablo Samuel CastroAAAI 2020 · 被引用 171 次
相关 Paper
- Denoised MDPs: Learning World Models Better Than the World ItselfTongzhou Wang, Simon S. Du, Antonio Torralba, Phillip Isola 等ICML 2022 · 被引用 63 次
- Learning Dynamic Attribute-factored World Models for Efficient Multi-object Reinforcement LearningFan Feng, Sara MagliacaneNeurIPS 2023 · 被引用 17 次
- Temporal Predictive Coding For Model-Based Planning In Latent SpaceTung D. Nguyen, Rui Shu, Tuan Pham, Hung Bui 等ICML 2021 · 被引用 65 次
- Building Minimal and Reusable Causal State Abstractions for Reinforcement LearningZizhao Wang, Caroline Wang, Xuesu Xiao, Yuke Zhu 等AAAI 2024 · 被引用 9 次
- Provable Rich Observation Reinforcement Learning with Combinatorial Latent StatesDipendra Misra, Qinghua Liu, Chi Jin, John LangfordICLR 2021 · 被引用 8 次
