Improving the Continuity of Goal-Achievement Ability via Policy Self-Regularization for Goal-Conditioned Reinforcement Learning
Xudong Gong, Sen Yang, Dawei Feng, Kele Xu, Bo Ding, Huaimin Wang, Yong Dou
摘要
Existing Goal-Conditioned Reinforcement Learning (GCRL) algorithms are built upon Hindsight Experience Replay (HER), which densifies rewards through hindsight replay and leverages historical goal-achieving information to construct a learning curriculum. However, when the task is characterized by a non-Markovian reward (NMR), whose computation depends on multiple steps of states and actions, HER can no longer densify rewards by treating a single encountered state as the hindsight goal. The lack of informative rewards hinders policy learning, resulting in rolling out failed trajectories. Consequently, the replay buffer is overwhelmed with failed trajectories, impeding the establishment of an applicable curriculum. To circumvent these limitations, we deviate from existing HER-based methods and propose an on-policy GCRL framework, GCPO, which is applicable to both multi-goal Markovian reward (MR) and NMR problems. GCPO consists of (1) Pre-training from Demonstrations, which pre-trains the policy to possess an initial goal-achieving capability, thereby diminishing the difficulty of subsequent online learning. (2) Online Self-Curriculum Learning, which first estimates the policy's goal-achieving capability based on historical evaluation information and then selects progressively challenging goals for learning based on its current capability. We evaluate GCPO on a challenging multi-goal long-horizon task: fixed-wing UAV velocity vector control. Experimental results demonstrate that GCPO is capable of effectively addressing both multi-goal MR and NMR problems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online VideosBowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga 等NeurIPS 2022 · 被引用 458 次
- Contrastive Learning as Goal-Conditioned Reinforcement LearningBenjamin Eysenbach, Tianjun Zhang, Sergey Levine, Ruslan SalakhutdinovNeurIPS 2022 · 被引用 331 次
- Learning to Reach Goals via Iterated Supervised LearningDibya Ghosh, Abhishek Gupta, Ashwin Reddy, Justin Fu 等ICLR 2021 · 被引用 222 次
- Explore, Discover and Learn: Unsupervised Discovery of State-Covering SkillsVictor Campos, Alexander Trott, Caiming Xiong, Richard Socher 等ICML 2020 · 被引用 178 次
- Planning Goals for ExplorationEdward S. Hu, Richard Chang, Oleh Rybkin, Dinesh JayaramanICLR 2023 · 被引用 152 次
相关 Paper
- Goal-Conditioned On-Policy Reinforcement LearningXudong Gong, Dawei Feng, Kele Xu, Bo Ding 等NeurIPS 2024 · 被引用 16 次
- First-Order Representation Languages for Goal-Conditioned RLSimon Ståhlberg, Hector GeffnerAAAI 2026 · 被引用 1 次
- VVC-Gym: A Fixed-Wing UAV Reinforcement Learning Environment for Multi-Goal Long-Horizon ProblemsXudong Gong, Dawei Feng, Kele Xu, Weijia Wang 等ICLR 2025
- Conservative Offline Goal-Conditioned Implicit V-LearningKaiqiang Ke, Qian Lin, Zongkai Liu, Shenghong He 等ICML 2025
- Pre-Trained Multi-Goal Transformers with Prompt Optimization for Efficient Online AdaptationHaoqi Yuan, Yuhui Fu, Feiyang Xie, Zongqing LuNeurIPS 2024 · 被引用 5 次
