Align-RUDDER: Learning From Few Demonstrations by Reward Redistribution
Vihang Patil, Markus Hofmarcher, Marius-Constantin Dinu, Matthias Dorfer, Patrick M. Blies, Johannes Brandstetter, José Antonio Arjona-Medina, Sepp Hochreiter
摘要
Reinforcement Learning algorithms require a large number of samples to solve complex tasks with sparse and delayed rewards. Complex tasks can often be hierarchically decomposed into sub-tasks. A step in the Q-function can be associated with solving a sub-task, where the expectation of the return increases. RUDDER has been introduced to identify these steps and then redistribute reward to them, thus immediately giving reward if sub-tasks are solved. Since the problem of delayed rewards is mitigated, learning is considerably sped up. However, for complex tasks, current exploration strategies as deployed in RUDDER struggle with discovering episodes with high rewards. Therefore, we assume that episodes with high rewards are given as demonstrations and do not have to be discovered by exploration. Typically the number of demonstrations is small and RUDDER's LSTM model as a deep learning method does not learn well. Hence, we introduce Align-RUDDER, which is RUDDER with two major modifications. First, Align-RUDDER assumes that episodes with high rewards are given as demonstrations, replacing RUDDER's safe exploration and lessons replay buffer. Second, we replace RUDDER's LSTM model by a profile model that is obtained from multiple sequence alignment of demonstrations. Profile models can be constructed from as few as two demonstrations as known from bioinformatics. Align-RUDDER inherits the concept of reward redistribution, which considerably reduces the delay of rewards, thus speeding up learning. Align-RUDDER outperforms competitors on complex artificial tasks with delayed reward and few demonstrations. On the MineCraft ObtainDiamond task, Align-RUDDER is able to mine a diamond, though not frequently. Github: this https URL, YouTube: this https URL
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Principle-Driven Self-Alignment of Language Models from Scratch with Minimal Human SupervisionZhiqing Sun, Yikang Shen, Qinhong Zhou, Hongxin Zhang 等NeurIPS 2023 · 被引用 463 次
- Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online VideosBowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga 等NeurIPS 2022 · 被引用 458 次
- Describe, Explain, Plan and Select: Interactive Planning with LLMs Enables Open-World Multi-Task AgentsZihao Wang, Shaofei Cai, Guanzhou Chen, Anji Liu 等NeurIPS 2023 · 被引用 178 次
- Do Embodied Agents Dream of Pixelated Sheep: Embodied Decision Making using Language Guided World ModellingKolby Nottingham, Prithviraj Ammanabrolu, Alane Suhr, Yejin Choi 等ICML 2023 · 被引用 110 次
- SALMON: Self-Alignment with Instructable Reward ModelsZhiqing Sun, Yikang Shen, Hongxin Zhang, Qinhong Zhou 等ICLR 2024 · 被引用 59 次
它引用的顶会 Paper5
- Hopfield Networks is All You NeedHubert Ramsauer, Bernhard Schäfl, Johannes Lehner, Philipp Seidl 等ICLR 2021 · 被引用 620 次
- Modern Hopfield Networks and Attention for Immune Repertoire ClassificationMichael Widrich, Bernhard Schäfl, Milena Pavlovic, Hubert Ramsauer 等NeurIPS 2020 · 被引用 152 次
- Reinforcement Learning from Imperfect Demonstrations under Soft Expert GuidanceMingxuan Jing, Xiaojian Ma, Wenbing Huang, Fuchun Sun 等AAAI 2020 · 被引用 70 次
- History Compression via Language Models in Reinforcement LearningFabian Paischer, Thomas Adler, Vihang Patil, Angela Bitto-Nemling 等ICML 2022 · 被引用 53 次
- Watch, Try, Learn: Meta-Learning from Demonstrations and RewardsAllan Zhou, Eric Jang, Daniel Kappler, Alexander Herzog 等ICLR 2020 · 被引用 53 次
相关 Paper
- DISCOVER: Automated Curricula for Sparse-Reward Reinforcement LearningLeander Diaz-Bone, Marco Bagatella, Jonas Hübotter, Andreas KrauseNeurIPS 2025 · 被引用 14 次
- Self-Paced Deep Reinforcement LearningPascal Klink, Carlo D'Eramo, Jan Peters, Joni PajarinenNeurIPS 2020 · 被引用 83 次
- Understanding the Complexity Gains of Single-Task RL with a CurriculumQiyang Li, Yuexiang Zhai, Yi Ma, Sergey LevineICML 2023 · 被引用 21 次
- Learning Guidance Rewards with Trajectory-space SmoothingTanmay Gangwani, Yuan Zhou, Jian PengNeurIPS 2020 · 被引用 46 次
- Off-Policy Reinforcement Learning with Delayed RewardsBeining Han, Zhizhou Ren, Zuofan Wu, Yuan Zhou 等ICML 2022 · 被引用 47 次
