Width-based Lookaheads with Learnt Base Policies and Heuristics Over the Atari-2600 Benchmark
Stefan O'Toole, Nir Lipovetzky, Miquel Ramírez, Adrian R. Pearce
Abstract
We propose new width-based planning and learning algorithms inspired from a careful analysis of the design decisions made by previous width-based planners. The algorithms are applied over the Atari-2600 games and our best performing algorithm, Novelty guided Critical Path Learning (N-CPL), outperforms the previously introduced width-based planning and learning algorithms -IW(1), -IW(1)+ and -HIW(n, 1). Furthermore, we present a taxonomy of the Atari-2600 games according to some of their defining characteristics. This analysis of the games provides further insight into the behaviour and performance of the algorithms introduced. Namely, for games with large branching factors, and games with sparse meaningful rewards, N-CPL outperforms -IW, -IW(1)+ and -HIW(n, 1).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on1
Related papers
- NovelD: A Simple yet Effective Exploration CriterionTianjun Zhang, Huazhe Xu, Xiaolong Wang, Yi Wu et al.NeurIPS 2021 · 106 citations
- On Bonus Based Exploration Methods In The Arcade Learning EnvironmentAdrien Ali Taïga, William Fedus, Marlos C. Machado, Aaron C. Courville et al.ICLR 2020 · 72 citations
- OptionZero: Planning with Learned OptionsPo-Wei Huang, Pei-Chiun Peng, Hung Guei, Ti-Rong WuICLR 2025
- Generalized Weighted Path Consistency for Mastering Atari GamesDengwei Zhao, Shikui Tu, Lei XuNeurIPS 2023 · 5 citations
- Distributional Meta-Gradient Reinforcement LearningHaiyan Yin, Shuicheng Yan, Zhongwen XuICLR 2023
