Episodic Policy Gradient Training
Hung Le, Majid Abdolshah, Thommen George Karimpanal, Kien Do, Dung Nguyen, Svetha Venkatesh
Abstract
We introduce a novel training procedure for policy gradient methods wherein episodic memory is used to optimize the hyperparameters of reinforcement learning algorithms on-the-fly. Unlike other hyperparameter searches, we formulate hyperparameter scheduling as a standard Markov Decision Process and use episodic memory to store the outcome of used hyperparameters and their training contexts. At any policy update step, the policy learner refers to the stored experiences, and adaptively reconfigures its learning algorithm with the new hyperparameters determined by the memory. This mechanism, dubbed as Episodic Policy Gradient Training (EPGT), enables an episodic learning process, and jointly learns the policy and the learning algorithm's hyperparameters within a single run. Experimental results on both continuous and discrete environments demonstrate the advantage of using the proposed method in boosting the performance of various policy gradient algorithms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Neural Episodic Control with State AbstractionZhuo Li, Derui Zhu, Yujing Hu, Xiaofei Xie et al.ICLR 2023 · 5 citations
- Stable Hadamard Memory: Revitalizing Memory-Augmented Agents for Reinforcement LearningHung Le, Dung Nguyen, Kien Do, Sunil Gupta et al.ICLR 2025
Builds on3
- Provably Efficient Online Hyperparameter Optimization with Population-Based BanditsJack Parker-Holder, Vu Nguyen, Stephen J. RobertsNeurIPS 2020 · 105 citations
- Self-Attentive Associative MemoryHung Le, Truyen Tran, Svetha VenkateshICML 2020 · 61 citations
- Model-Based Episodic Memory Induces Dynamic Hybrid ControlsHung Le, Thommen George Karimpanal, Majid Abdolshah, Truyen Tran et al.NeurIPS 2021 · 25 citations
Related papers
- Tuning Mixed Input Hyperparameters on the Fly for Efficient Population Based AutoRLJack Parker-Holder, Vu Nguyen, Shaan Desai, Stephen J. RobertsNeurIPS 2021 · 22 citations
- Scalable One-Pass Optimisation of High-Dimensional Weight-Update Hyperparameters by Implicit DifferentiationRoss M. Clarke, Elre Talea Oldewage, José Miguel Hernández-LobatoICLR 2022 · 9 citations
- Experience-Evolving Multi-Turn Tool-Use Agent with Hybrid Episodic–Procedural MemorySijia Li, Yuchen Huang, Zifan LIU, Zijian LI et al.ICML 2026 · 3 citations
- Steady State Analysis of Episodic Reinforcement LearningBojun HuangNeurIPS 2020 · 29 citations
- Phasic Policy GradientKarl Cobbe, Jacob Hilton, Oleg Klimov, John SchulmanICML 2021 · 191 citations
