Memory Based Trajectory-conditioned Policies for Learning from Sparse Rewards
Yijie Guo, Jongwook Choi, Marcin Moczulski, Shengyu Feng, Samy Bengio, Mohammad Norouzi, Honglak Lee
摘要
Reinforcement learning with sparse rewards is challenging because an agent can rarely obtain non-zero rewards and hence, gradient-based optimization of parameterized policies can be incremental and slow. Recent work demonstrated that using a memory buffer of previous successful trajectories can result in more effective policies. However, existing methods may overly exploit past successful experiences, which can encourage the agent to adopt sub-optimal and myopic behaviors. In this work, instead of focusing on good experiences with limited diversity, we propose to learn a trajectory-conditioned policy to follow and expand diverse past trajectories from a memory buffer. Our method allows the agent to reach diverse regions in the state space and improve upon the past trajectories to reach new states. We empirically show that our approach significantly outperforms count-based exploration methods (parametric approach) and self-imitation learning (parametric approach with non-parametric memory) on various complex tasks with local optima. In particular, without using expert demonstrations or resetting to arbitrary states, we achieve the state-of-the-art scores under five billion number of frames, on challenging Atari games such as Montezuma's Revenge and Pitfall.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Planning Goals for ExplorationEdward S. Hu, Richard Chang, Oleh Rybkin, Dinesh JayaramanICLR 2023 · 被引用 152 次
- A Robust and Opponent-Aware League Training Method for StarCraft IIRuozi Huang, Xipeng Wu, Hongsheng Yu, Zhong Fan 等NeurIPS 2023 · 被引用 13 次
- GUIDE: Real-Time Human-Shaped AgentsLingyu Zhang, Zhengran Ji, Nicholas R. Waytowich, Boyuan ChenNeurIPS 2024 · 被引用 9 次
- Learning to Execute: Efficient Learning of Universal Plan-Conditioned Policies in RoboticsIngmar Schubert, Danny Driess, Ozgur S. Oguz, Marc ToussaintNeurIPS 2021 · 被引用 2 次
- Reinforcement Learning Enhanced Muti-hop Reasoning for Temporal Knowledge Question AnsweringWuzhenghong Wen, Chao Xue, Su Pan, Yuwei Sun 等AAAI 2026
它引用的顶会 Paper4
- Model Based Reinforcement Learning for AtariLukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski 等ICLR 2020 · 被引用 969 次
- Never Give Up: Learning Directed Exploration StrategiesAdrià Puigdomènech Badia, Pablo Sprechmann, Alex Vitvitskyi, Zhaohan Daniel Guo 等ICLR 2020 · 被引用 349 次
- Skew-Fit: State-Covering Self-Supervised Reinforcement LearningVitchyr Pong, Murtaza Dalal, Steven Lin, Ashvin Nair 等ICML 2020 · 被引用 303 次
- Behaviour Suite for Reinforcement LearningIan Osband, Yotam Doron, Matteo Hessel, John Aslanides 等ICLR 2020 · 被引用 204 次
相关 Paper
- Learning to Constrain Policy Optimization with Virtual Trust RegionHung Le, Thommen Karimpanal George, Majid Abdolshah, Dung Nguyen 等NeurIPS 2022 · 被引用 5 次
- Locally Persistent Exploration in Continuous Control Tasks with Sparse RewardsSusan Amin, Maziar Gomrokchi, Hossein Aboutalebi, Harsh Satija 等ICML 2021 · 被引用 17 次
- Learning World Models for Unconstrained Goal NavigationYuanlin Duan, Wensen Mao, He ZhuNeurIPS 2024 · 被引用 11 次
- Monte Carlo Augmented Actor-Critic for Sparse Reward Deep Reinforcement Learning from Suboptimal DemonstrationsAlbert Wilcox, Ashwin Balakrishna, Jules Dedieu, Wyame Benslimane 等NeurIPS 2022 · 被引用 28 次
- Generalizable Episodic Memory for Deep Reinforcement LearningHao Hu, Jianing Ye, Guangxiang Zhu, Zhizhou Ren 等ICML 2021 · 被引用 44 次
