A Long N-step Surrogate Stage Reward for Deep Reinforcement Learning
Junmin Zhong, Ruofan Wu, Jennie Si
摘要
We introduce a new stage reward estimator named the long N -step surrogate stage (LNSS) reward for deep reinforcement learning (RL). It aims at mitigating the high variance problem, which has shown impeding successful convergence of learning, hurting task performance, and hindering applications of deep RL in continuous control problems. In this paper we show that LNSS, which utilizes a long reward trajectory of future steps, provides consistent performance improvement measured by average reward, convergence speed, learning success rate, and variance reduction in Q values and rewards. Our evaluations are based on a variety of environments in DeepMind Control Suite and OpenAI Gym by using LNSS piggybacked on baseline deep RL algorithms such as DDPG, D4PG, and TD3. We show that LNSS reward has enabled good results that have been challenging to obtain by deep RL previously. Our analysis also shows that LNSS exponentially reduces the upper bound on the variances of Q values from respective single-step methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- State Chrono Representation for Enhancing Generalization in Reinforcement LearningJianda Chen, Wen Zheng Terence Ng, Zichen Chen, Sinno Jialin Pan 等NeurIPS 2024 · 被引用 5 次
- Reinforcement Learning Control of a Physical Robot Device for Assisted Human Walking without a SimulatorJunmin Zhong, Emiliano Quiñones Yumbla, Seyed Yousef Soltanian, Ruofan Wu 等ICML 2025
它引用的顶会 Paper3
- Reinforcement Learning with Perturbed RewardsJingkang Wang, Yang Liu, Bo LiAAAI 2020 · 被引用 161 次
- Is High Variance Unavoidable in RL? A Case Study in Continuous ControlJohan Bjorck, Carla P. Gomes, Kilian Q. WeinbergerICLR 2022 · 被引用 36 次
- Human-Robotic Prosthesis as Collaborating Agents for Symmetrical WalkingRuofan Wu, Junmin Zhong, Brent Wallace, Xiang Gao 等NeurIPS 2022 · 被引用 16 次
相关 Paper
- Deterministic Value-Policy GradientsQingpeng Cai, Ling Pan, Pingzhong TangAAAI 2020 · 被引用 1 次
- Efficient Continuous Control with Double Actors and Regularized CriticsJiafei Lyu, Xiaoteng Ma, Jiangpeng Yan, Xiu LiAAAI 2022 · 被引用 69 次
- Learning Guidance Rewards with Trajectory-space SmoothingTanmay Gangwani, Yuan Zhou, Jian PengNeurIPS 2020 · 被引用 46 次
- Taylor TD-learningMichele Garibbo, Maxime Robeyns, Laurence AitchisonNeurIPS 2023
- Foresee then Evaluate: Decomposing Value Estimation with Latent Future PredictionHongyao Tang, Zhaopeng Meng, Guangyong Chen, Pengfei Chen 等AAAI 2021 · 被引用 5 次
