DRIBO: Robust Deep Reinforcement Learning via Multi-View Information Bottleneck
Jiameng Fan, Wenchao Li
Abstract
Deep reinforcement learning (DRL) agents are often sensitive to visual changes that were unseen in their training environments. To address this problem, we leverage the sequential nature of RL to learn robust representations that encode only task-relevant information from observations based on the unsupervised multi-view setting. Specifically, we introduce a novel contrastive version of the Multi-View Information Bottleneck (MIB) objective for temporal data. We train RL agents from pixels with this auxiliary objective to learn robust representations that can compress away task-irrelevant information and are predictive of task-relevant dynamics. This approach enables us to train high-performance policies that are robust to visual distractions and can generalize well to unseen environments. We demonstrate that our approach can achieve SOTA performance on a diverse set of visual control tasks in the DeepMind Control Suite when the background is replaced with natural videos. In addition, we show that our approach outperforms well-established baselines for generalization to unseen environments on the Procgen benchmark. Our code is open-sourced and available at https://github. com/BU-DEPEND-Lab/DRIBO.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e2c4086e-d10c-4b21-bedf-ae975551e1b6Cited by top-tier papers14
- Look where you look! Saliency-guided Q-networks for generalization in visual Reinforcement LearningDavid Bertoin, Adil Zouitine, Mehdi Zouitine, Emmanuel RachelsonNeurIPS 2022 · 67 citations
- PLASTIC: Improving Input and Label Plasticity for Sample Efficient Reinforcement LearningHojoon Lee, Hanseul Cho, Hyunseung Kim, Daehoon Gwak et al.NeurIPS 2023 · 50 citations
- On the Importance of Exploration for Generalization in Reinforcement LearningYiding Jiang, J. Zico Kolter, Roberta RaileanuNeurIPS 2023 · 48 citations
- Revisiting Plasticity in Visual Reinforcement Learning: Data, Modules and Training StagesGuozheng Ma, Lu Li, Sen Zhang, Zixuan Liu et al.ICLR 2024 · 32 citations
- Selective Visual Representations Improve Convergence and Generalization for Embodied AIAinaz Eftekhar, Kuo-Hao Zeng, Jiafei Duan, Ali Farhadi et al.ICLR 2024 · 28 citations
Builds on15
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Data-Efficient Image Recognition with Contrastive Predictive CodingOlivier J. HénaffICML 2020 · 1,553 citations
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 1,261 citations
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 1,170 citations
Related papers
- DreamerPro: Reconstruction-Free Model-Based Reinforcement Learning with Prototypical RepresentationsFei Deng, Ingook Jang, Sungjin AhnICML 2022 · 83 citations
- Dream to Generalize: Zero-Shot Model-Based Reinforcement Learning for Unseen Visual DistractionsJeongsoo Ha, Kyungsoo Kim, Yusung KimAAAI 2023 · 10 citations
- Temporal Disentanglement of Representations for Improved Generalisation in Reinforcement LearningMhairi Dunion, Trevor McInroe, Kevin Sebastian Luck, Josiah P. Hanna et al.ICLR 2023 · 4 citations
- Learning Invariant Representations for Reinforcement Learning without ReconstructionAmy Zhang, Rowan Thomas McAllister, Roberto Calandra, Yarin Gal et al.ICLR 2021 · 77 citations
- Temporal Predictive Coding For Model-Based Planning In Latent SpaceTung D. Nguyen, Rui Shu, Tuan Pham, Hung Bui et al.ICML 2021 · 65 citations
