Learning Robust Representations with Long-Term Information for Generalization in Visual Reinforcement Learning
Rui Yang, Jie Wang, Qijie Peng, Ruibo Guo, Guoping Wu, Bin Li
Abstract
Reinforcement learning algorithms have achieved remarkable success in acquiring behavioral skills directly from pixel inputs. However, their application in real-world scenarios presents challenges due to their sensitivity to visual distractions (e.g., changes in viewpoint and light). A key factor contributing to this challenge is that the learned representations often suffer from overfitting taskirrelevant information. By comparing several representation learning methods, we find that the key to alleviating overfitting in representation learning is to choose proper prediction targets. Motivated by our comparison, we propose a novel representation learning approach-namely, reward sequence prediction (RSP)-that uses reward sequences or their transforms (e.g., discrete time Fourier transform) as prediction targets. RSP can efficiently learn robust representations as reward sequences rarely contain task-irrelevant information while providing a large number of supervised signals to accelerate representation learning. An appealing feature is that RSP makes no assumption about the type of distractions and thus can improve performance even when multiple types of distractions exist. We evaluate our approach in Distracting Control Suite. Experiments show that our method achieves state-of-the-art sample efficiency and generalization ability in tasks with distractions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 403e5671-a417-438e-b95f-ccd6d7a6e822Cited by top-tier papers4
- Multi-Agent Reinforcement Learning with Communication-Constrained PriorsGuang Yang, Tianpei Yang, Jingwen Qiao, Yanqing Wu et al.NeurIPS 2025 · 9 citations
- Diffusion Guided Adaptive Augmentation for Generalization in Visual Reinforcement LearningJeong Woon Lee, Hyoseok HwangICCV 2025 · 3 citations
- Task-Aware Exploration via a Predictive Bisimulation MetricDayang Liang, Ruihan LIU, Lipeng Wan, Yunlong Liu et al.ICML 2026 · 1 citation
- Fourier Guided Adaptive Adversarial Augmentation for Generalization in Visual Reinforcement LearningJeong Woon Lee, Hyoseok HwangAAAI 2025
Builds on25
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 1,261 citations
- Domain Generalization with MixStyleKaiyang Zhou, Yongxin Yang, Yu Qiao, Tao XiangICLR 2021 · 986 citations
- Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from PixelsDenis Yarats, Ilya Kostrikov, Rob FergusICLR 2021 · 911 citations
- Reinforcement Learning with Augmented DataMichael Laskin, Kimin Lee, Adam Stooke, Lerrel Pinto et al.NeurIPS 2020 · 833 citations
Related papers
- Learning Task-relevant Representations for Generalization via Characteristic Functions of Reward Sequence DistributionsRui Yang, Jie Wang, Zijie Geng, Mingxuan Ye et al.KDD 2022 · 13 citations
- Robust Representation Learning by Clustering with Bisimulation Metrics for Visual Reinforcement Learning with DistractionsQiyuan Liu, Qi Zhou, Rui Yang, Jie WangAAAI 2023 · 22 citations
- Task-Induced Representation LearningJun Yamada, Karl Pertsch, Anisha Gunjal, Joseph J. LimICLR 2022 · 15 citations
- Denoised MDPs: Learning World Models Better Than the World ItselfTongzhou Wang, Simon S. Du, Antonio Torralba, Phillip Isola et al.ICML 2022 · 63 citations
- Policy-Independent Behavioral Metric-Based Representation for Deep Reinforcement LearningWeijian Liao, Zongzhang Zhang, Yang YuAAAI 2023 · 7 citations
