Denoised MDPs: Learning World Models Better Than the World Itself
Tongzhou Wang, Simon S. Du, Antonio Torralba, Phillip Isola, Amy Zhang, Yuandong Tian
摘要
The ability to separate signal from noise, and reason with clean abstractions, is critical to intelligence. With this ability, humans can efficiently perform real world tasks without considering all possible nuisance factors. How can artificial agents do the same? What kind of information can agents safely discard as noises? In this work, we categorize information out in the wild into four types based on controllability and relation with reward, and formulate useful information as that which is both controllable and reward-relevant. This framework clarifies the kinds information removed by various prior work on representation learning in reinforcement learning (RL), and leads to our proposed approach of learning a Denoised MDP that explicitly factors out certain noise distractors. Extensive experiments on variants of DeepMind Control Suite and RoboDesk demonstrate superior performance of our denoised world model over using raw observations alone, and over prior works, across policy optimization control tasks as well as the non-control task of joint position regression.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper31
- Bridging State and History Representations: Understanding Self-Predictive RLTianwei Ni, Benjamin Eysenbach, Erfan Seyedsalehi, Michel Ma 等ICLR 2024 · 被引用 50 次
- Language Control Diffusion: Efficiently Scaling through Space, Time, and TasksEdwin Zhang, Yujie Lu, Shinda Huang, William Yang Wang 等ICLR 2024 · 被引用 34 次
- CCIL: Continuity-Based Data Augmentation for Corrective Imitation LearningLiyiming Ke, Yunchu Zhang, Abhay Deshpande, Siddhartha S. Srinivasa 等ICLR 2024 · 被引用 33 次
- Learning World Models with Identifiable FactorizationYuren Liu, Biwei Huang, Zhengmao Zhu, Hong-Long Tian 等NeurIPS 2023 · 被引用 32 次
- RePo: Resilient Model-Based Reinforcement Learning by Regularizing Posterior PredictabilityChuning Zhu, Max Simchowitz, Siri Gadipudi, Abhishek GuptaNeurIPS 2023 · 被引用 24 次
它引用的顶会 Paper10
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 被引用 2,360 次
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 被引用 1,261 次
- Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from PixelsDenis Yarats, Ilya Kostrikov, Rob FergusICLR 2021 · 被引用 911 次
- Reinforcement Learning with Augmented DataMichael Laskin, Kimin Lee, Adam Stooke, Lerrel Pinto 等NeurIPS 2020 · 被引用 833 次
相关 Paper
- Learning Robust Representations with Long-Term Information for Generalization in Visual Reinforcement LearningRui Yang, Jie Wang, Qijie Peng, Ruibo Guo 等ICLR 2025
- Leveraging Conditional Dependence for Efficient World Model DenoisingShaowei Zhang, Jiahan Cao, Dian Cheng, Xunlan Zhou 等NeurIPS 2025
- DreamerPro: Reconstruction-Free Model-Based Reinforcement Learning with Prototypical RepresentationsFei Deng, Ingook Jang, Sungjin AhnICML 2022 · 被引用 83 次
- Task-Induced Representation LearningJun Yamada, Karl Pertsch, Anisha Gunjal, Joseph J. LimICLR 2022 · 被引用 15 次
- Policy-Independent Behavioral Metric-Based Representation for Deep Reinforcement LearningWeijian Liao, Zongzhang Zhang, Yang YuAAAI 2023 · 被引用 7 次
