A Simple Framework for Generalization in Visual RL under Dynamic Scene Perturbations
Wonil Song, Hyesong Choi, Kwanghoon Sohn, Dongbo Min
摘要
In the rapidly evolving domain of vision-based deep reinforcement learning (RL), a pivotal challenge is to achieve generalization capability to dynamic environmental changes reflected in visual observations. Our work delves into the intricacies of this problem, identifying two key issues that appear in previous approaches for visual RL generalization: (i) imbalanced saliency and (ii) observational overfitting. Imbalanced saliency is a phenomenon where an RL agent disproportionately identifies salient features across consecutive frames in a frame stack. Observational overfitting occurs when the agent focuses on certain background regions rather than task-relevant objects. To address these challenges, we present a simple yet effective framework for generalization in visual RL (SimGRL) under dynamic scene perturbations. First, to mitigate the imbalanced saliency problem, we introduce an architectural modification to the image encoder to stack frames at the feature level rather than the image level. Simultaneously, to alleviate the observational overfitting problem, we propose a novel technique called shifted random overlay augmentation, which is specifically designed to learn robust representations capable of effectively handling dynamic visual scenes. Extensive experiments demonstrate the superior generalization capability of SimGRL, achieving state-of-the-art performance in benchmarks including the DeepMind Control Suite. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Focus-Then-Reuse: Fast Adaptation in Visual Perturbation EnvironmentsJiahui Wang, Chao Chen, Jiacheng Xu, Zongzhang Zhang 等NeurIPS 2025 · 被引用 1 次
- PQDA: Policy-Aligned Q-Consistency Meets Decoupled Augmentation for Generalizable Visual RLYun Zhou, Yuqiang Wu, Chunyu TanAAAI 2026
- TSTM: Temporal Segmentation for Task-relevant Mask in Visual Reinforcement Learning GeneralizationWeicheng Du, Wenjia Meng, Zhengzhe Zhang, Yilong Yin 等CVPR 2026
- Disentangled World Models: Learning to Transfer Semantic Knowledge from Distracting Videos for Reinforcement LearningQi Wang, Zhipeng Zhang, Baao Xie, Xin Jin 等ICCV 2025
它引用的顶会 Paper18
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 被引用 1,261 次
- Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from PixelsDenis Yarats, Ilya Kostrikov, Rob FergusICLR 2021 · 被引用 911 次
- Reinforcement Learning with Augmented DataMichael Laskin, Kimin Lee, Adam Stooke, Lerrel Pinto 等NeurIPS 2020 · 被引用 833 次
- Improving Sample Efficiency in Model-Free Reinforcement Learning from ImagesDenis Yarats, Amy Zhang, Ilya Kostrikov, Brandon Amos 等AAAI 2021 · 被引用 506 次
相关 Paper
- Focus On What Matters: Separated Models For Visual-Based RL GeneralizationDi Zhang, Bowen Lv, Hai Zhang, Feifan Yang 等NeurIPS 2024 · 被引用 14 次
- Spectrum Random Masking for Generalization in Image-based Reinforcement LearningYangru Huang, Peixi Peng, Yifan Zhao, Guangyao Chen 等NeurIPS 2022 · 被引用 33 次
- Diffusion Guided Adaptive Augmentation for Generalization in Visual Reinforcement LearningJeong Woon Lee, Hyoseok HwangICCV 2025 · 被引用 3 次
- Simoun: Synergizing Interactive Motion-appearance Understanding for Vision-based Reinforcement LearningYangru Huang, Peixi Peng, Yifan Zhao, Yunpeng Zhai 等ICCV 2023 · 被引用 3 次
- Environment Agnostic Representation for Visual Reinforcement learningHyesong Choi, Hunsang Lee, Seongwon Jeong, Dongbo MinICCV 2023 · 被引用 17 次
