Width-based Lookaheads with Learnt Base Policies and Heuristics Over the Atari-2600 Benchmark
Stefan O'Toole, Nir Lipovetzky, Miquel Ramírez, Adrian R. Pearce
摘要
We propose new width-based planning and learning algorithms inspired from a careful analysis of the design decisions made by previous width-based planners. The algorithms are applied over the Atari-2600 games and our best performing algorithm, Novelty guided Critical Path Learning (N-CPL), outperforms the previously introduced width-based planning and learning algorithms -IW(1), -IW(1)+ and -HIW(n, 1). Furthermore, we present a taxonomy of the Atari-2600 games according to some of their defining characteristics. This analysis of the games provides further insight into the behaviour and performance of the algorithms introduced. Namely, for games with large branching factors, and games with sparse meaningful rewards, N-CPL outperforms -IW, -IW(1)+ and -HIW(n, 1).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper1
相关 Paper
- NovelD: A Simple yet Effective Exploration CriterionTianjun Zhang, Huazhe Xu, Xiaolong Wang, Yi Wu 等NeurIPS 2021 · 被引用 106 次
- On Bonus Based Exploration Methods In The Arcade Learning EnvironmentAdrien Ali Taïga, William Fedus, Marlos C. Machado, Aaron C. Courville 等ICLR 2020 · 被引用 72 次
- OptionZero: Planning with Learned OptionsPo-Wei Huang, Pei-Chiun Peng, Hung Guei, Ti-Rong WuICLR 2025
- Generalized Weighted Path Consistency for Mastering Atari GamesDengwei Zhao, Shikui Tu, Lei XuNeurIPS 2023 · 被引用 5 次
- Distributional Meta-Gradient Reinforcement LearningHaiyan Yin, Shuicheng Yan, Zhongwen XuICLR 2023
