The Variational Bandwidth Bottleneck: Stochastic Evaluation on an Information Budget
Anirudh Goyal, Yoshua Bengio, Matthew M. Botvinick, Sergey Levine
摘要
In many applications, it is desirable to extract only the relevant information from complex input data, which involves making a decision about which input features are relevant. The information bottleneck method formalizes this as an information-theoretic optimization problem by maintaining an optimal tradeoff between compression (throwing away irrelevant input information), and predicting the target. In many problem settings, including the reinforcement learning problems we consider in this work, we might prefer to compress only part of the input. This is typically the case when we have a standard conditioning input, such as a state observation, and a ``privileged'' input, which might correspond to the goal of a task, the output of a costly planning algorithm, or communication with another agent. In such cases, we might prefer to compress the privileged input, either to achieve better generalization (e.g., with respect to goals) or to minimize access to costly information (e.g., in the case of communication). Practical implementations of the information bottleneck based on variational inference require access to the privileged input in order to compute the bottleneck variable, so although they perform compression, this compression operation itself needs unrestricted, lossless access. In this work, we propose the variational bandwidth bottleneck, which decides for each example on the estimated value of the privileged information before seeing it, i.e., only based on the standard input, and then accordingly chooses stochastically, whether to access the privileged input or not. We formulate a tractable approximation to this framework and demonstrate in a series of reinforcement learning experiments that it can improve generalization and reduce access to computationally costly information.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Predictive Information Accelerates Learning in RLKuang-Huei Lee, Ian Fischer, Anthony Z. Liu, Yijie Guo 等NeurIPS 2020 · 被引用 82 次
- Minimum Description Length ControlTed Moskovitz, Ta-Chu Kao, Maneesh Sahani, Matt M. BotvinickICLR 2023 · 被引用 76 次
- Retrieval-Augmented Reinforcement LearningAnirudh Goyal, Abram L. Friesen, Andrea Banino, Theophane Weber 等ICML 2022 · 被引用 69 次
- CoMic: Complementary Task Learning & Mimicry for Reusable SkillsLeonard Hasenclever, Fabio Pardo, Raia Hadsell, Nicolas Heess 等ICML 2020 · 被引用 56 次
- Understanding the Complexity Gains of Single-Task RL with a CurriculumQiyang Li, Yuexiang Zhai, Yi Ma, Sergey LevineICML 2023 · 被引用 21 次
相关 Paper
- Drop-Bottleneck: Learning Discrete Compressed Representation for Noise-Robust ExplorationJaekyeom Kim, Minjung Kim, Dongyeon Woo, Gunhee KimICLR 2021 · 被引用 20 次
- Explaining A Black-box By Using A Deep Variational Information Bottleneck ApproachSeo-Jin Bang, Pengtao Xie, Heewook Lee, Wei Wu 等AAAI 2021 · 被引用 33 次
- Adaptive Discrete Communication Bottlenecks with Dynamic Vector Quantization for Heterogeneous Representational CoarsenessDianbo Liu, Alex Lamb, Xu Ji, Pascal Tikeng Notsawo Jr. 等AAAI 2023 · 被引用 19 次
- Information Retention via Learning Supplemental FeaturesZhipeng Xie, Yahe LiICLR 2024 · 被引用 1 次
- Learning Efficient Multi-agent Communication: An Information Bottleneck ApproachRundong Wang, Xu He, Runsheng Yu, Wei Qiu 等ICML 2020 · 被引用 133 次
