What Effects the Generalization in Visual Reinforcement Learning: Policy Consistency with Truncated Return Prediction
Shuo Wang, Zhihao Wu, Xiaobo Hu, Jinwen Wang, Youfang Lin, Kai Lv
Abstract
In visual Reinforcement Learning (RL), the challenge of generalization to new environments is paramount. This study pioneers a theoretical analysis of visual RL generalization, establishing an upper bound on the generalization objective, encompassing policy divergence and Bellman error components. Motivated by this analysis, we propose maintaining the cross-domain consistency for each policy in the policy space, which can reduce the divergence of the learned policy during the test. In practice, we introduce the Truncated Return Prediction (TRP) task, promoting cross-domain policy consistency by predicting truncated returns of historical trajectories. Moreover, we also propose a Transformer-based predictor for this auxiliary task. Extensive experiments on Deep-Mind Control Suite and Robotic Manipulation tasks demonstrate that TRP achieves state-of-the-art generalization performance. We further demonstrate that TRP outperforms previous methods in terms of sample efficiency during training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Seek Commonality but Preserve Differences: Dissected Dynamics Modeling for Multi-modal Visual RLYangru Huang, Peixi Peng, Yifan Zhao, Guangyao Chen et al.NeurIPS 2024 · 3 citations
- Out-of-Distribution Generalization with a SPARC: Racing 100 Unseen Vehicles with a Single PolicyBram Grooten, Patrick MacAlpine, Kaushik Subramanian, Peter Stone et al.AAAI 2026 · 2 citations
- From Pixels to Temporal Correlations: Learning Informative Representations for Reinforcement Learning Pre-trainingJinwen Wang, Youfang Lin, Xiaobo Hu, Siyu Yang et al.ACM MM 2025 · 1 citation
- TSTM: Temporal Segmentation for Task-relevant Mask in Visual Reinforcement Learning GeneralizationWeicheng Du, Wenjia Meng, Zhengzhe Zhang, Yilong Yin et al.CVPR 2026
Builds on16
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 1,261 citations
- Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from PixelsDenis Yarats, Ilya Kostrikov, Rob FergusICLR 2021 · 911 citations
- Reinforcement Learning with Augmented DataMichael Laskin, Kimin Lee, Adam Stooke, Lerrel Pinto et al.NeurIPS 2020 · 833 citations
- Data-Efficient Reinforcement Learning with Self-Predictive RepresentationsMax Schwarzer, Ankesh Anand, Rishab Goel, R. Devon Hjelm et al.ICLR 2021 · 399 citations
Related papers
- Prompt-based Visual Alignment for Zero-shot Policy TransferHaihan Gao, Rui Zhang, Qi Yi, Hantao Yao et al.ICML 2024 · 1 citation
- Return-Critic: Bridging Goal Discrepancy for Efficient Visual Reinforcement LearningRuyi Lu, Xuesong Wang, Hengrui Zhang, Yuhu ChengICML 2026
- Saliency-Guided Representation with Consistency Policy Learning for Visual Unsupervised Reinforcement LearningJingbo Sun, Qichao Zhang, Songjun Tu, Xing Fang et al.CVPR 2026 · 1 citation
- Learning Task-relevant Representations for Generalization via Characteristic Functions of Reward Sequence DistributionsRui Yang, Jie Wang, Zijie Geng, Mingxuan Ye et al.KDD 2022 · 13 citations
- Generalizing Consistency Policy to Visual RL with Prioritized Proximal Experience RegularizationHaoran Li, Zhennan Jiang, Yuhui Chen, Dongbin ZhaoNeurIPS 2024 · 16 citations
