Visual Reinforcement Learning with Residual Action
Zhenxian Liu, Peixi Peng, Yonghong Tian
Abstract
Learning control policy from continuous action space by visual observations is a fundamental and challenging task in reinforcement learning (RL). An essential problem is how to accurately map the high-dimensional images to the optimal actions by the policy network. Traditional decision-making modules output actions solely based on the current observation, while the distributions of optimal actions are dependent on specific tasks and cannot be known priorly, which increases the learning difficulty. To make the learning easier, we analyze the action characteristics in several control tasks, and propose Reinforcement Learning with Residual Action (ResAct) to explicitly model the adjustments of actions based on the differences between adjacent observations, rather than learning actions directly from observations. The method just redefines the output of the policy network, and doesn’t introduce any prior assumption to constrain or simplify the vanilla control problem. Extensive experiments on DeepMind Control Suite and CARLA demonstrate that the method could improve different RL baselines significantly, and achieve state-of-the-art performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Dejavu: Towards Experience Feedback Learning for Embodied IntelligenceShaokai Wu, Yanbiao Ji, Qiuchang Li, Zhiyi Zhang et al.CVPR 2026 · 3 citations
- Multi-timescale Reinforcement Learning by Value ReconstructionZhan Su, Peixi Peng, Xinyu Hu, Cong Li et al.ICML 2026
- Resolving the Stability-Plasticity Dilemma in Reinforcement Learning via Complementary Continual CriticsBo Sun, Peixi Peng, Guang Tan, Haoran Xu et al.CVPR 2026
- TSTM: Temporal Segmentation for Task-relevant Mask in Visual Reinforcement Learning GeneralizationWeicheng Du, Wenjia Meng, Zhengzhe Zhang, Yilong Yin et al.CVPR 2026
- Return-Critic: Bridging Goal Discrepancy for Efficient Visual Reinforcement LearningRuyi Lu, Xuesong Wang, Hengrui Zhang, Yuhu ChengICML 2026
Builds on19
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 1,261 citations
- Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from PixelsDenis Yarats, Ilya Kostrikov, Rob FergusICLR 2021 · 911 citations
- Reinforcement Learning with Augmented DataMichael Laskin, Kimin Lee, Adam Stooke, Lerrel Pinto et al.NeurIPS 2020 · 833 citations
- Stochastic Latent Actor-Critic: Deep Reinforcement Learning with a Latent Variable ModelAlex X. Lee, Anusha Nagabandi, Pieter Abbeel, Sergey LevineNeurIPS 2020 · 437 citations
Related papers
- Look where you look! Saliency-guided Q-networks for generalization in visual Reinforcement LearningDavid Bertoin, Adil Zouitine, Mehdi Zouitine, Emmanuel RachelsonNeurIPS 2022 · 67 citations
- Residual Q-Learning: Offline and Online Policy Customization without ValueChenran Li, Chen Tang, Haruki Nishimura, Jean Mercat et al.NeurIPS 2023 · 15 citations
- Policy Gradient With Serial Markov Chain ReasoningEdoardo Cetin, Oya ÇeliktutanNeurIPS 2022 · 4 citations
- Imitation by Predicting ObservationsAndrew Jaegle, Yury Sulsky, Arun Ahuja, Jake Bruce et al.ICML 2021 · 16 citations
- Learning Task-relevant Representations for Generalization via Characteristic Functions of Reward Sequence DistributionsRui Yang, Jie Wang, Zijie Geng, Mingxuan Ye et al.KDD 2022 · 13 citations
