A Long N-step Surrogate Stage Reward for Deep Reinforcement Learning
Junmin Zhong, Ruofan Wu, Jennie Si
Abstract
We introduce a new stage reward estimator named the long N -step surrogate stage (LNSS) reward for deep reinforcement learning (RL). It aims at mitigating the high variance problem, which has shown impeding successful convergence of learning, hurting task performance, and hindering applications of deep RL in continuous control problems. In this paper we show that LNSS, which utilizes a long reward trajectory of future steps, provides consistent performance improvement measured by average reward, convergence speed, learning success rate, and variance reduction in Q values and rewards. Our evaluations are based on a variety of environments in DeepMind Control Suite and OpenAI Gym by using LNSS piggybacked on baseline deep RL algorithms such as DDPG, D4PG, and TD3. We show that LNSS reward has enabled good results that have been challenging to obtain by deep RL previously. Our analysis also shows that LNSS exponentially reduces the upper bound on the variances of Q values from respective single-step methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 30082781-35bd-4d19-91f9-afbf06843f85Cited by top-tier papers2
- State Chrono Representation for Enhancing Generalization in Reinforcement LearningJianda Chen, Wen Zheng Terence Ng, Zichen Chen, Sinno Jialin Pan et al.NeurIPS 2024 · 5 citations
- Reinforcement Learning Control of a Physical Robot Device for Assisted Human Walking without a SimulatorJunmin Zhong, Emiliano Quiñones Yumbla, Seyed Yousef Soltanian, Ruofan Wu et al.ICML 2025
Builds on3
- Reinforcement Learning with Perturbed RewardsJingkang Wang, Yang Liu, Bo LiAAAI 2020 · 161 citations
- Is High Variance Unavoidable in RL? A Case Study in Continuous ControlJohan Bjorck, Carla P. Gomes, Kilian Q. WeinbergerICLR 2022 · 36 citations
- Human-Robotic Prosthesis as Collaborating Agents for Symmetrical WalkingRuofan Wu, Junmin Zhong, Brent Wallace, Xiang Gao et al.NeurIPS 2022 · 16 citations
Related papers
- Deterministic Value-Policy GradientsQingpeng Cai, Ling Pan, Pingzhong TangAAAI 2020 · 1 citation
- Efficient Continuous Control with Double Actors and Regularized CriticsJiafei Lyu, Xiaoteng Ma, Jiangpeng Yan, Xiu LiAAAI 2022 · 69 citations
- Learning Guidance Rewards with Trajectory-space SmoothingTanmay Gangwani, Yuan Zhou, Jian PengNeurIPS 2020 · 46 citations
- Taylor TD-learningMichele Garibbo, Maxime Robeyns, Laurence AitchisonNeurIPS 2023
- Foresee then Evaluate: Decomposing Value Estimation with Latent Future PredictionHongyao Tang, Zhaopeng Meng, Guangyong Chen, Pengfei Chen et al.AAAI 2021 · 5 citations
