Memory Based Trajectory-conditioned Policies for Learning from Sparse Rewards
Yijie Guo, Jongwook Choi, Marcin Moczulski, Shengyu Feng, Samy Bengio, Mohammad Norouzi, Honglak Lee
Abstract
Reinforcement learning with sparse rewards is challenging because an agent can rarely obtain non-zero rewards and hence, gradient-based optimization of parameterized policies can be incremental and slow. Recent work demonstrated that using a memory buffer of previous successful trajectories can result in more effective policies. However, existing methods may overly exploit past successful experiences, which can encourage the agent to adopt sub-optimal and myopic behaviors. In this work, instead of focusing on good experiences with limited diversity, we propose to learn a trajectory-conditioned policy to follow and expand diverse past trajectories from a memory buffer. Our method allows the agent to reach diverse regions in the state space and improve upon the past trajectories to reach new states. We empirically show that our approach significantly outperforms count-based exploration methods (parametric approach) and self-imitation learning (parametric approach with non-parametric memory) on various complex tasks with local optima. In particular, without using expert demonstrations or resetting to arbitrary states, we achieve the state-of-the-art scores under five billion number of frames, on challenging Atari games such as Montezuma's Revenge and Pitfall.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b9048917-30d8-4909-8eb0-e28aab3abd1aCited by top-tier papers5
- Planning Goals for ExplorationEdward S. Hu, Richard Chang, Oleh Rybkin, Dinesh JayaramanICLR 2023 · 152 citations
- A Robust and Opponent-Aware League Training Method for StarCraft IIRuozi Huang, Xipeng Wu, Hongsheng Yu, Zhong Fan et al.NeurIPS 2023 · 13 citations
- GUIDE: Real-Time Human-Shaped AgentsLingyu Zhang, Zhengran Ji, Nicholas R. Waytowich, Boyuan ChenNeurIPS 2024 · 9 citations
- Learning to Execute: Efficient Learning of Universal Plan-Conditioned Policies in RoboticsIngmar Schubert, Danny Driess, Ozgur S. Oguz, Marc ToussaintNeurIPS 2021 · 2 citations
- Reinforcement Learning Enhanced Muti-hop Reasoning for Temporal Knowledge Question AnsweringWuzhenghong Wen, Chao Xue, Su Pan, Yuwei Sun et al.AAAI 2026
Builds on4
- Model Based Reinforcement Learning for AtariLukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski et al.ICLR 2020 · 969 citations
- Never Give Up: Learning Directed Exploration StrategiesAdrià Puigdomènech Badia, Pablo Sprechmann, Alex Vitvitskyi, Zhaohan Daniel Guo et al.ICLR 2020 · 349 citations
- Skew-Fit: State-Covering Self-Supervised Reinforcement LearningVitchyr Pong, Murtaza Dalal, Steven Lin, Ashvin Nair et al.ICML 2020 · 303 citations
- Behaviour Suite for Reinforcement LearningIan Osband, Yotam Doron, Matteo Hessel, John Aslanides et al.ICLR 2020 · 204 citations
Related papers
- Learning to Constrain Policy Optimization with Virtual Trust RegionHung Le, Thommen Karimpanal George, Majid Abdolshah, Dung Nguyen et al.NeurIPS 2022 · 5 citations
- Locally Persistent Exploration in Continuous Control Tasks with Sparse RewardsSusan Amin, Maziar Gomrokchi, Hossein Aboutalebi, Harsh Satija et al.ICML 2021 · 17 citations
- Learning World Models for Unconstrained Goal NavigationYuanlin Duan, Wensen Mao, He ZhuNeurIPS 2024 · 11 citations
- Monte Carlo Augmented Actor-Critic for Sparse Reward Deep Reinforcement Learning from Suboptimal DemonstrationsAlbert Wilcox, Ashwin Balakrishna, Jules Dedieu, Wyame Benslimane et al.NeurIPS 2022 · 28 citations
- Generalizable Episodic Memory for Deep Reinforcement LearningHao Hu, Jianing Ye, Guangxiang Zhu, Zhizhou Ren et al.ICML 2021 · 44 citations
