Robust Predictable Control
Ben Eysenbach, Ruslan Salakhutdinov, Sergey Levine
摘要
Many of the challenges facing today's reinforcement learning (RL) algorithms, such as robustness, generalization, transfer, and computational efficiency are closely related to compression. Prior work has convincingly argued why minimizing information is useful in the supervised learning setting, but standard RL algorithms lack an explicit mechanism for compression. The RL setting is unique because (1) its sequential nature allows an agent to use past information to avoid looking at future observations and (2) the agent can optimize its behavior to prefer states where decision making requires few bits. We take advantage of these properties to propose a method (RPC) for learning simple policies. This method brings together ideas from information bottlenecks, model-based RL, and bits-back coding into a simple and theoretically-justified algorithm. Our method jointly optimizes a latent-space model and policy to be self-consistent, such that the policy avoids states where the model is inaccurate. We demonstrate that our method achieves much tighter compression than prior methods, achieving up to 5× higher reward than a standard information bottleneck. We also demonstrate that our method learns policies that are more robust and generalize better to new tasks. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Dynamics Generalisation in Reinforcement Learning via Adaptive Context-Aware PoliciesMichael Beukman, Devon Jarvis, Richard Klein, Steven James 等NeurIPS 2023 · 被引用 29 次
- Predictable MDP Abstraction for Unsupervised Model-Based RLSeohong Park, Sergey LevineICML 2023 · 被引用 11 次
- Maximum Total Correlation Reinforcement LearningBang You, Puze Liu, Huaping Liu, Jan Peters 等ICML 2025
- The Value of Sensory Information to a RobotArjun Krishna, Edward S. Hu, Dinesh JayaramanICLR 2025
- Watch Less, Do More: Implicit Skill Discovery for Video-Conditioned PolicyJiangxing Wang, Zongqing LuICLR 2025
它引用的顶会 Paper6
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 被引用 1,261 次
- Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from PixelsDenis Yarats, Ilya Kostrikov, Rob FergusICLR 2021 · 被引用 911 次
- Stochastic Latent Actor-Critic: Deep Reinforcement Learning with a Latent Variable ModelAlex X. Lee, Anusha Nagabandi, Pieter Abbeel, Sergey LevineNeurIPS 2020 · 被引用 437 次
- Adversarial Robustness vs. Model Compression, or Both?Shaokai Ye, Xue Lin, Kaidi Xu, Sijia Liu 等ICCV 2019 · 被引用 180 次
- Learning Efficient Multi-agent Communication: An Information Bottleneck ApproachRundong Wang, Xu He, Runsheng Yu, Wei Qiu 等ICML 2020 · 被引用 133 次
相关 Paper
- Reinforcement Learning with Simple Sequence PriorsTankred Saanum, Noémi Élteto, Peter Dayan, Marcel Binz 等NeurIPS 2023 · 被引用 18 次
- From Parameters to Behaviors: Unsupervised Compression of the Policy SpaceDavide Tenedini, Riccardo Zamboni, Mirco Mutti, Marcello RestelliICLR 2026 · 被引用 4 次
- Drop-Bottleneck: Learning Discrete Compressed Representation for Noise-Robust ExplorationJaekyeom Kim, Minjung Kim, Dongyeon Woo, Gunhee KimICLR 2021 · 被引用 20 次
- DRIBO: Robust Deep Reinforcement Learning via Multi-View Information BottleneckJiameng Fan, Wenchao LiICML 2022 · 被引用 49 次
- The Variational Bandwidth Bottleneck: Stochastic Evaluation on an Information BudgetAnirudh Goyal, Yoshua Bengio, Matthew M. Botvinick, Sergey LevineICLR 2020 · 被引用 26 次
