Episodic Policy Gradient Training
Hung Le, Majid Abdolshah, Thommen George Karimpanal, Kien Do, Dung Nguyen, Svetha Venkatesh
摘要
We introduce a novel training procedure for policy gradient methods wherein episodic memory is used to optimize the hyperparameters of reinforcement learning algorithms on-the-fly. Unlike other hyperparameter searches, we formulate hyperparameter scheduling as a standard Markov Decision Process and use episodic memory to store the outcome of used hyperparameters and their training contexts. At any policy update step, the policy learner refers to the stored experiences, and adaptively reconfigures its learning algorithm with the new hyperparameters determined by the memory. This mechanism, dubbed as Episodic Policy Gradient Training (EPGT), enables an episodic learning process, and jointly learns the policy and the learning algorithm's hyperparameters within a single run. Experimental results on both continuous and discrete environments demonstrate the advantage of using the proposed method in boosting the performance of various policy gradient algorithms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Neural Episodic Control with State AbstractionZhuo Li, Derui Zhu, Yujing Hu, Xiaofei Xie 等ICLR 2023 · 被引用 5 次
- Stable Hadamard Memory: Revitalizing Memory-Augmented Agents for Reinforcement LearningHung Le, Dung Nguyen, Kien Do, Sunil Gupta 等ICLR 2025
它引用的顶会 Paper3
- Provably Efficient Online Hyperparameter Optimization with Population-Based BanditsJack Parker-Holder, Vu Nguyen, Stephen J. RobertsNeurIPS 2020 · 被引用 105 次
- Self-Attentive Associative MemoryHung Le, Truyen Tran, Svetha VenkateshICML 2020 · 被引用 61 次
- Model-Based Episodic Memory Induces Dynamic Hybrid ControlsHung Le, Thommen George Karimpanal, Majid Abdolshah, Truyen Tran 等NeurIPS 2021 · 被引用 25 次
相关 Paper
- Tuning Mixed Input Hyperparameters on the Fly for Efficient Population Based AutoRLJack Parker-Holder, Vu Nguyen, Shaan Desai, Stephen J. RobertsNeurIPS 2021 · 被引用 22 次
- Scalable One-Pass Optimisation of High-Dimensional Weight-Update Hyperparameters by Implicit DifferentiationRoss M. Clarke, Elre Talea Oldewage, José Miguel Hernández-LobatoICLR 2022 · 被引用 9 次
- Experience-Evolving Multi-Turn Tool-Use Agent with Hybrid Episodic–Procedural MemorySijia Li, Yuchen Huang, Zifan LIU, Zijian LI 等ICML 2026 · 被引用 3 次
- Steady State Analysis of Episodic Reinforcement LearningBojun HuangNeurIPS 2020 · 被引用 29 次
- Phasic Policy GradientKarl Cobbe, Jacob Hilton, Oleg Klimov, John SchulmanICML 2021 · 被引用 191 次
